Latent Stationary Overlap and Minimax Off-policy Evaluation in Finite POMDPs
Abstract
This paper characterizes the observable-data minimax risk for stationary off-policy evaluation in finite partially observed Markov decision processes. The data are one behavior-policy trajectory of length , the trajectory horizon, with bounded rewards and binary actions; policies are known functions of the observed state; the joint observed–latent Markov state satisfies stationary start, sequential ignorability, one-step policy overlap , where is the log-overlap scale, geometric total-variation contraction , where is the mixing scale, and latent stationary overlap , where is the stationary-overlap constant and and are the target and behavior stationary joint-state laws. For sufficiently large , and cardinality-uniformly over finite observed and latent alphabets, the minimax squared-error risk is comparable, uniformly over , to where is the overlap-distance radius, the normalized distance of the stationary density-ratio bound from unit overlap, and is the rate exponent set by the mixing and log-overlap scales, up to constants depending only on and . For every fixed , the surface has the Hu–Wager partial-history exponent ; the parametric term changes the rate on shrinking overlap-distance scales . A clipped partial-history importance weighting (PHIW) estimator with radius-calibrated history depth attains this surface when the depth is calibrated from . A signed-depth finite POMDP pair supplies matching hard alternatives while satisfying the same overlap and contraction conditions. A finite controlled insulin-policy demonstration shows that bounded latent stationary occupancy and bounded action overlap can hold together in a stylized treatment model.
Introduction
Off-policy evaluation asks how well the value of a target policy can be learned from data generated by a different behavior policy. In longitudinal causal inference this question is tied to the g-formula, propensity-score balancing, marginal structural models, and dynamic treatment regimes (Robins, 1986; Rosenbaum et al., 1983; Robins et al., 2000; Murphy, 2003). In reinforcement learning and stochastic control it is the policy-evaluation problem for Markov decision processes and partially observed Markov decision processes (Smallwood et al., 1973; Kaelbling et al., 1998; Sutton et al., 2018). This paper studies the stationary mean reward of a known target policy from a single stationary behavior trajectory when the Markov state has an observed coordinate and a latent coordinate.
The fully observed benchmark is by now well developed. Importance weighting, doubly robust methods, and stationary density-ratio methods use the observed Markov state to replace full trajectory likelihood ratios by more stable value-function or occupancy-ratio objects (Precup et al., 2000; Jiang et al., 2016; Thomas et al., 2016; Liu et al., 2018; Nachum et al., 2019; Uehara et al., 2020; Kallus et al., 2020). Econometric work on adaptive and longitudinal treatment settings gives related average-reward and policy-learning theory (Liao et al., 2021; Liao et al., 2022). These analyses identify the role of overlap in observed states and actions when the state available to the analyst is rich enough to be Markov.
Partial observation changes the statistical object. Hu et al. (2023) show that, in finite-state POMDPs with known policies, bounded rewards, one-step policy overlap, and geometric mixing, partial-history importance weighting can evaluate a target policy at a polynomial mean-squared-error rate governed by the mixing and overlap constants. The estimator uses a recent observable history: a longer history reduces latent-state bias, while multiplying more policy ratios increases variance. The corresponding minimax lower bound establishes the same partial-history exponent within the stated bounded-reward, binary-action finite-state model class.
This paper’s main message is that bounding the latent stationary density ratio at any fixed radius preserves the partially observed Hu–Wager exponent . The overlap-distance radius changes the rate in the local layer where shrinks on the scale. The model class in Definition 10 consists of finite observed and latent alphabets, known memoryless policies on , a stationary behavior start, bounded rewards, sequential ignorability, one-step policy overlap with log-overlap scale , uniform total-variation contraction with mixing scale , and latent stationary overlap for the target and behavior stationary joint-state laws. The estimand in Definition 12 is the stationary target-policy mean reward over the joint observed–latent state, and denotes expectation under the observed law induced by model . The minimax comparison ranges over finite alphabets with constants depending only on fixed after the threshold in Theorem 1; estimators observe the observed trajectory, its observed alphabet size, and the policy pair.
Assumption interpretation
Latent stationary overlap is plausible when changing from the behavior policy to the target policy changes the long-run frequency of each joint observed–latent regime by a bounded factor. An analyst would justify it through design knowledge, mechanistic modeling, sensitivity analysis, or external evidence about how the target policy shifts latent regimes. Because the bound is imposed on the joint state, it is stronger than overlap for the observed state alone and carries information about unobserved occupancy. Sequential ignorability and known memoryless binary policies match randomized or logged-policy settings in which treatment probabilities are functions of the observed state. Uniform total-variation contraction is a Doeblin-type mixing condition for both policies; it supplies the stationary laws, the bias decay in the PHIW upper bound, and the signed-depth membership argument. Bounded rewards keep squared-error risk on a common scale. Cardinality-uniform finite alphabets mean that the risk constants are stable as the observed and latent alphabets vary.
Limitations
Latent stationary overlap cannot be checked from the observed trajectory alone. Uniform Doeblin-type contraction can fail in slowly mixing clinical dynamics, in which case the PHIW bias term and the signed-depth lower-bound calibration require model-specific replacement. Non-stationary starts require burn-in or explicit initial-distribution terms beyond the stationary-start statements used here. The lower-bound alternatives use latent alphabet size of order , so a sharp frontier for a fixed hidden alphabet remains an open question. The radius-calibrated depth attains the displayed frontier using , equivalently , and data-driven selection for an unknown overlap radius, including the possible cost of logarithmic adaptation, remains open. The local path is a mathematical local-asymptotic slice of the uniform surface.
The main result is the all-radius minimax surface in Theorem 1. For fixed and , define the exponent and the overlap-distance radius . Uniformly over every , the cardinality-uniform observable-data minimax squared-error risk is comparable to with constants and a large-sample threshold depending only on and . The radius-calibrated partial-history estimator in Definition 20 and Definition 16 attains the upper bound; its depth is computed from , so the overlap radius is a tuning input supplied to the estimator rather than a quantity it learns from the data. The lower bound combines the radius-explicit signed-depth construction with a uniform parametric floor.
The formula has three immediate interpretations. At the unit-overlap boundary , the radius is zero, the target and behavior stationary laws coincide, and Theorem 3 gives the minimax scale. For every fixed , the radius is constant and Theorem 2 gives the partially observed rate . Along local paths with , Theorem 4 gives the transition up to constants, with an elbow at the scale . Below that scale, immediate weighting reaches the parametric order; above it, the radius-calibrated PHIW depth used in the upper bound grows logarithmically with the effective radius.
The lower-bound alternatives are finite signed-depth POMDPs. Their hidden state records a depth coordinate and a sign coordinate; the target policy forces the rare action, while the behavior policy takes that action with probability . The reward signal appears at terminal depth, and the stationary target value changes sign across the two alternatives. Lemma 18 places these hard alternatives in the latent-overlap model class, with amplitude comparable to the radius through . The observable divergence bounds in Lemmas 31 and 35 pass through the signed-depth factor , and Lemma 40 supplies the parametric floor. Together these statements match the bias–variance difficulty faced by PHIW at the radius-calibrated depth.
The paper also includes a finite controlled insulin-policy demonstration. The construction in Definition 24 specifies glucose, diet, activity, and lagged treatment coordinates in a 360-state controlled Markov model designed to make the assumptions checkable by finite arithmetic. The target policy treats above a glucose threshold and the behavior policy randomizes treatment. Theorem 5 verifies finite state count, one-half contraction, bounded one-step policy overlap, finite latent stationary overlap with the sufficient bound , and a strictly negative interval arithmetic bound for the immediate-weighting diagnostic. This value of gives , so the depth prescribed at this radius by Definition 20 is once , and zero on the complementary branch. Because is within of one, this depth differs from the fixed-radius balanced depth of Definition 17 only by a bounded rounding adjustment at large . Since the reward is a function of observed glucose, the diagnostic compares behavior-stationary and target-stationary reward averages in this stylized finite treatment model.
The appendix closes with notes on the formal results, presentation-level definitions, and cited external inputs.
The paper proceeds as follows. The related-work section places the result among dynamic treatment regimes, off-policy evaluation, and partially observed models. The setup section defines the finite POMDP class, the stationary target value, minimax risk, the partial-history estimator, and the signed-depth family. The main-results section states the uniform all-radius theorem first and then its fixed-overlap, unit-overlap, and shrinking-radius regimes. The insulin-policy section gives the finite treatment-policy adaptation and its interval-certified diagnostic. The discussion interprets the risk surface and records limitations. The appendix collects the auxiliary constructions, covariance bounds, signed-depth likelihood calculations, radius-sensitive risk bounds, and parametric lower-bound ingredients.
Related work
The paper sits at the intersection of dynamic treatment regimes, off-policy evaluation, and partially observed Markov decision processes. Sequential causal estimands trace back to the potential-outcome and g-formula traditions of Robins (1986), Rosenbaum et al. (1983), and Robins et al. (2000), with dynamic regimes developed in econometrics, biostatistics, and reinforcement learning by work such as Murphy (2003). In that tradition, the central object is a target-policy value identified from data generated by a different behavior policy under sequential ignorability and overlap. The present analysis keeps that causal interpretation, while specializing to finite-state partially observed Markov decision processes with a stationary target value and a latent stationary overlap condition.
A second line of work studies off-policy evaluation in Markov decision processes. Classical dynamic-programming and POMDP foundations appear in Smallwood et al. (1973), Kaelbling et al. (1998), and Sutton et al. (2018). For fully observed MDPs, importance weighting, density-ratio estimation, and doubly robust methods provide estimators whose behavior depends on policy ratios, state-occupancy ratios, and mixing properties (Precup et al., 2000; Jiang et al., 2016; Thomas et al., 2016; Liu et al., 2018; Nachum et al., 2019; Uehara et al., 2020; Kallus et al., 2020). Recent econometric work extends these ideas to adaptive and longitudinal settings, including Liao et al. (2021) and Liao et al. (2022). These methods supply the fully observed benchmark: when the state observed by the analyst is the Markov state, occupancy-ratio control can stabilize long-horizon evaluation. The minimax comparisons in Theorems 2, 4, and 1 instead quantify the cost of evaluating a stationary target policy when the overlap restriction is imposed on the latent stationary law.
The closest partially observed benchmark is Hu et al. (2023). Their finite-state POMDP analysis uses sequential ignorability, known policies, bounded rewards, one-step policy overlap, and geometric mixing to study partial-history estimators and the minimax partial-history rate. The new ingredient here is the latent stationary overlap radius on the joint hidden–observed state. For each fixed , Theorem 2 keeps the Hu–Wager exponent ; the fixed-overlap upper bound uses the balanced PHIW depth, and the lower bound uses signed-depth alternatives satisfying the same latent-overlap radius. A separate construction is needed for the converse, and the reason is a short calculation. In the chain built in the proof of Hu et al. (2023), the hidden coordinate is a depth ladder ; one action advances the chain one rung with probability , the other resets it to , the target policy always advances, and the behavior policy advances with probability , where is their one-step log-overlap scale. We record the consequence as a short calculation of our own, read off their construction rather than quoted from a numbered result: reaching the terminal rung requires consecutive advances, so the two stationary masses at that rung are and their ratio is , exponential in the depth. Because their rate-calibrated choice takes , that ratio eventually exceeds any fixed : their hard instances leave for every fixed radius, so the fixed-radius converse here is obtained from the signed-depth family of Definition 18 instead, whose reset step is perturbed by precisely so that the stationary ratio stays inside the radius.
A closely related fully observed comparison is Mehrabi et al. (2025). Their state is the Markov state, and they weaken the usual requirement that the target-to-behavior stationary density ratio be uniformly bounded to a tail condition on that ratio: Mehrabi et al. (2025) asks only that the -th power of the stationary ratio have a bounded tail under the behavior stationary law. In the square-integrable regime their truncated doubly robust estimator is -consistent and asymptotically normal with the semiparametric efficiency variance (Mehrabi et al., 2025). Four differences keep that result and the present one from being two readings of one condition. Their overlap norm is a tail bound on a ratio of observed-state stationary laws, whereas Assumption 7 is a pointwise bound on a ratio of joint observed–latent stationary laws. Their estimator is given the Markov state, whereas the estimators admitted by Definition 14 never see the latent coordinate. Their state space is general and the bounded-ratio assumption they relax is credible mainly on bounded state spaces, whereas the class here is uniform over all finite observed and latent cardinalities. Finally, their inferential target is a limit distribution for one estimator, whereas the target here is a worst-case squared-error rate over the whole class. Theorem 1 shows the resulting contrast: fixed latent stationary overlap with any fixed retains the partial-history rate , while the parametric term affects the frontier on shrinking scales . Theorems 3 and 4 record the endpoint and local slices of the same uniform surface.
A related POMDP literature recovers target values through proxy, bridge, coverage, or observability structure. Proxy and bridge approaches such as Tennenholtz et al. (2020), Bennett et al. (2021), Miao et al. (2022), Shi et al. (2022), Uehara et al. (2023), and Kuang et al. (2025) impose relationships that make observed variables informative about latent states or bridge functions. Zhang et al. (2024) give coverage conditions for future-dependent value functions, and Zhang et al. (2025) separates tractable and hard regimes for history-dependent policies. The present minimax analysis works in the finite POMDP experiment class with sequential ignorability, bounded rewards, contraction, one-step policy overlap, and latent stationary overlap, and quantifies the partial-history evaluation rate implied by those conditions.
The insulin-policy example connects the abstract finite-state model to dynamic treatment settings where the observed record is a coarse clinical history and the physiological state contains latent components. Prior work on diabetes management and treatment-policy learning, including Maahs et al. (2012), Luckett et al. (2020), and related reinforcement-learning formulations such as Nahum-Shani et al. (2018), motivates finite controlled models with clinically interpretable state variables and actions. The adaptation in Theorem 5 provides a finite controlled Markov model on which four of the paper’s structural constants are certified by finite arithmetic — state count, contraction rate, one-step policy overlap, and latent stationary overlap — together with a certified nonzero immediate-weighting bias, linking the mechanism behind the theory to a stylized treatment-policy evaluation problem.
Setup and estimands
We observe a single trajectory from a finite partially observed Markov decision process. Throughout, denotes the trajectory horizon, is total variation distance, and all index ranges are the finite ranges displayed in the definitions below. Generic constants may depend on the primitive radius and mixing parameters explicitly named in a result, while the minimax statements are uniform over finite observed and latent cardinalities. Asymptotic notation is used only for the horizon regime governed by the fixed parameters in the displayed model classes.
The finite alphabets are the observed state space , the latent state space , and the binary action set . The joint state has observed coordinate and latent coordinate ; the action and reward are and . The behavior policy generates the data, the target policy defines the counterfactual policy value, and the one-step law governs rewards and successor states.
For a real parameter , define the policy-overlap constant by
⊢ LeanThe positive log-overlap scale parameterizes the policy-overlap constant . Larger values allow the target policy to place relatively more mass on actions that are less common under the behavior policy.
For natural numbers , a full trajectory of length is a pair Here is the observed–latent joint state at state time , with observed coordinate and latent coordinate . At transition time , is the Boolean action and is the reward; when actions are written numerically, is identified with and with .
⊢ LeanThe full trajectory records both the observed and latent coordinates. The econometrician observes only the observed coordinate, action, and reward, but the structural restrictions are most compactly stated on the finite observed–latent state process.
The next group of conditions specifies the controlled Markov experiment and the way actions enter the observed-data law.
For each ,
⊢ LeanAssumption 1 is the time-homogeneous POMDP condition used in Model 3 of Hu et al. (2023). It makes the current joint state and action sufficient for the next reward and state transition, so the latent coordinate carries the unobserved predictive information.
For , a memoryless policy on an observed-state alphabet of size is a function When the action alphabet is written as , we identify with and with . This is a carrier only: no nonnegativity and no normalization is imposed here. The probability-vector restrictions on the behavior and target carriers are imposed later, on in Assumption 2 and on in Assumption 5.
⊢ LeanMemoryless policies depend on the observed coordinate rather than on the latent state. This matches the information set of the policy comparison: both the behavior and target policies are revealed functions of observed state and action. The object just defined is only the carrier for such a function; that the behavior carrier and the target carrier are genuine action distributions is a restriction imposed by the assumptions below, on in Assumption 2 and on in Assumption 5.
Identify the numeric action alphabet with the Boolean action set by and . The behavior carrier is a probability vector at each observed state, For every , the joint law of the state-and-epoch history strictly before , the current joint state , and the current action factors as the law of that history and current joint state followed by the behavior kernel evaluated at the current observed state: Equivalently, for every and every ,
⊢ LeanAssumption 2 is the observed-state sequential ignorability and randomization condition in Hu et al. (2023). Conditional on the current observed state, action assignment follows the behavior policy even when the latent coordinate affects future rewards and states.
For a raw experiment with one-step kernel , a memoryless policy , and joint states and , define the policy-induced transition operator by
⊢ LeanThe policy-induced transition operator is the reward-marginalized Markov kernel on joint states obtained after averaging actions under policy . It is the transition law used to define stationarity and mixing under both the behavior and target policies.
For natural numbers and , the finite observed–latent joint-state alphabet is
⊢ LeanThe joint alphabet is finite but otherwise unrestricted. This cardinality-uniform formulation lets the risk statements range over both observed and latent state complexities.
For a memoryless policy , let be the policy-induced Markov kernel on the joint state space . Define to be the stationary joint-state law satisfying We write for the behavior-policy stationary joint-state law and for the target-policy stationary joint-state law.
The stationary law provides the population distribution attached to a fixed memoryless policy. In particular, describes the stationary behavior experiment, while supplies the stationary target distribution in the value estimand.
We now impose the remaining restrictions that define the population experiment analyzed below.
The initial joint state has the behavior stationary law :
⊢ LeanAssumption 3 is the behavior-stationary initialization condition in Hu et al. (2023). It aligns the observed trajectory with the stationary behavior distribution from the first period.
For each ,
⊢ LeanAssumption 4 is the uniformly bounded rewards condition in Hu et al. (2023). The common scale fixes the squared-error normalization used in the minimax risk.
The target carrier is a probability vector at each observed state, and for every the policy-overlap constant satisfies
⊢ LeanAssumption 5 is the uniform one-step policy overlap condition in Hu et al. (2023). It bounds observable action likelihood ratios state by state, which is the primitive overlap control for partial-history weighting, and it also records that the target policy is a genuine action distribution at every observed state, which is what makes the one-step ratios average to one.
For any real , define the contraction parameter by
⊢ LeanThe positive mixing scale determines the contraction parameter . Smaller corresponds to faster geometric forgetting in total variation distance.
For every , the contraction parameter and the policy-induced transition operator satisfy
⊢ LeanAssumption 6 is the uniform geometric total-variation contraction condition in Hu et al. (2023). It supplies a common mixing rate for the behavior and target policy kernels on the joint state space.
For the overlap constant and experiment , the latent stationary overlap condition holds when and, writing and for the stationary laws of the target-policy and behavior-policy joint-state transition kernels, respectively,
⊢ LeanAssumption 7 is the strong stationary distributional overlap condition imposed on the unobserved joint state. The stationary density-ratio bound relates the target and behavior stationary laws over the full finite state space, including the latent coordinate.
The observed-data object is a trajectory of observed coordinates, actions, and rewards. The next definitions collect the seven restrictions into the model class and define the value and risks used in the main results.
For a horizon and observed-state cardinality , define the observable trajectory space by Thus an element of records, at each epoch , the observed coordinate , the binary action , and the real reward . Equivalently, it is the sequence .
⊢ LeanThe observable trajectory space is the statistical sample space for estimators. It keeps the latent path out of the estimator input while retaining the observed policy information and .
For an observable-data POMDP experiment with , , policy-induced kernels , stationary laws , , and , write for the following seven restrictions, respectively: as in Assumption 1; as in Assumption 2; as in Assumption 3; as in Assumption 4; as in Assumption 5; as in Assumption 6; and as in Assumption 7.
Definition 9 gives short predicate names for the seven restrictions. These names are bookkeeping devices for the class definition and preserve the direct link between the compact model notation and the assumptions just stated.
For and , is the class of all observable-data POMDP experiments with the following defining properties:
(Horizon and alphabets.) The trajectory horizon satisfies , the observed and latent alphabets and are finite, and .
(Parameter scales.) The mixing and log-overlap parameters satisfy and .
(Raw experiment.) The tuple has a reward-transition kernel , behavior policy , target policy , and probability law for the observed trajectory generated from the joint state coordinates .
(Seven restrictions.) The seven restrictions in Assumptions 1, 2, 3, 4, 5, 6, and 7 hold, with policy-overlap factor and contraction factor .
Equivalently, The class ranges over all finite observed and latent cardinalities and all revealed behavior-target policy pairs satisfying these conditions.
⊢ LeanThe model class is the population domain for the minimax analysis. Its parameters separately control the time horizon, mixing scale, one-step action overlap, and latent stationary overlap.
For any raw finite-state experiment with target policy and one-step law , and for every joint state , define the target-policy reward regression by
⊢ LeanThe target-policy reward regression averages the one-period reward under the target action distribution at a fixed joint state. It is the reward component of the stationary target value.
For the joint state and the target-policy reward regression ,
⊢ LeanThe estimand is the stationary mean reward obtained by running the target policy and evaluating rewards under its stationary joint-state law. It is the population target for observable-data off-policy evaluation.
Let be a model in the latent-overlap class of Definition 10, whose observed trajectory is and let denote the probability law of induced by under the behavior policy .
For any observable-data estimator , its estimator-specific worst-case squared risk over is where is the revealed observed-alphabet size of . In particular, is this quantity evaluated at the clipped partial-history importance-weighted estimator with balanced history depth .
The estimator-specific risk in Definition 13 evaluates a fixed observable-data rule uniformly over the model class. This quantity is used for the partial-history estimator bound, while the minimax risk below optimizes over all admissible observable-data estimators.
For a trajectory horizon and parameters , the cardinality-uniform observable-data minimax squared-error risk is Here is the finite observed- and latent-state model class in Definition 10, with unrestricted finite observed alphabet size and unrestricted finite latent alphabet size, and is the stationary target-policy mean reward in Definition 12. The positive-horizon requirement is part of membership in , not a separate restriction on the supremum; at the degenerate horizon the class is empty and the display is not used. The infimum ranges over all observable-data estimators that take as inputs only , , , and the observed trajectory view , are measurable in for each , and are uniformly valued in .
⊢ LeanThe minimax risk measures the best achievable squared-error performance from the observed trajectory and the revealed policies. The uniformity over finite cardinalities makes the rate a statement about the structural restrictions rather than a fixed finite state dimension.
The final definitions in this section introduce the concrete estimator and the lower-bound family used in the main results.
For behavior and target policy functions , the observable one-step policy ratio is
⊢ LeanThe observable one-step policy ratio compares target and behavior action probabilities along the observed path. The zero-cell convention is compatible with the domination condition in Assumption 5.
For , define where is the observable one-step policy ratio with the zero-cell convention. The clipped partial-history importance-weighted estimator is
⊢ LeanThe score weights the reward by the recent observable policy-ratio history, and the estimator averages these scores with clipping to the reward scale. The depth is the tuning parameter that trades the contraction of latent-history bias against the variance from multiplying policy ratios.
For the contraction parameter and policy-overlap constant , the balanced fixed-radius history depth is
⊢ LeanThe fixed-depth rule balances the two primitives and at horizon . It is the estimator depth used in the fixed-radius upper bound.
For a trajectory length , parameters , a terminal depth , and an alternative index , the signed-depth family is defined under It is the raw POMDP experiment in with one observed state and hidden states obtained by embedding the following finite-reward model. The observed coordinate is constant, and the hidden coordinate is written Writing for the mixing value determined by , for the policy factor determined by , , and , the behavior and target policies are The initial hidden-state law is the probability measure obtained from the weights The transition kernel is the probability measure on reward-symbol and next-hidden-state pairs whose real weight at , with , is the product of the reward weight and the hidden-transition weight specified as follows. A reset sign has weight If , then only transitions with receive this reset-sign weight. If , transitions with receive weight , while transitions with receive weight and all other hidden transitions receive weight zero. The finite reward symbols are mapped to real rewards by and , and the trajectory law is the finite chronological law generated by this initial distribution, behavior policy, and kernel, then embedded into the unrestricted real-reward raw experiment.
⊢ LeanThe signed-depth two-point sub-experiment in Definition 18 is the finite construction used for the lower bounds. Its terminal depth , alternative sign , latent depth coordinate , latent sign coordinate , reset-sign perturbation , reward-amplitude constant , and truncated geometric mass make the overlap and contraction parameters explicit inside a finite POMDP.
Main results
The preceding section defined the observed-data experiment class, the stationary target value, the minimax risk, and the partial-history importance-weighted estimators. This section gives the risk surface as the latent stationary overlap radius varies. The organizing parameters are the normalized distance from unit overlap and the radius-calibrated history depth, where the radius is supplied to the estimator as a tuning input.
For , define
⊢ LeanThe radius is zero at equality of the two stationary laws and increases with the allowed stationary density-ratio bound. It is the scale on which the hidden-state part of the risk is measured.
For and , define
⊢ LeanThe depth shortens the history when the radius is close to unit overlap and lengthens it when the latent stationary shift is large enough to make hidden-state memory relevant. The threshold is the parametric boundary layer in the theorem below.
For the all-radius theorem it is convenient to index models directly and write their target values without carrying the kernel and policy notation in each display.
For , , and , let be the finite observed-word law through horizon generated by the signed-depth construction of Definition 18. Its observed state is constant, its latent state is , its behavior and target policies satisfy and , and its reward alphabet is with The initial latent law is , and the resulting finite observation consists of the action-reward word , equivalently the finite observed trajectory with constant observed-state coordinate.
The next theorem is the master characterization. Its constants depend on and , while the comparison holds uniformly over every radius and over the finite observed and hidden cardinalities allowed by Definition 10.
For an indexed model , write and for its joint reward-transition kernel and target policy. The target value of is with the stationary target-policy causal value from Definition 12.
Theorem 1 gives a single observable estimator rule — one formula, applied at every radius — together with a matching minimax surface. The rule is not radius-free: the depth it selects is computed from the supplied value of , so different radii give different depths and different estimators. The expression for has a parametric term and a hidden-state term scaled by . The radius-calibrated rule in Definition 20 implements this balance directly: in the boundary layer it uses depth zero, and away from that layer it increases the observable history length according to the effective radius .
The lower-bound calculations use the finite observed law induced by the signed-depth construction through horizon . Naming that law separates the observable experiment used in the converse from the latent construction that generates it.
For the balanced history depth , define the unclipped partial-history average by Here on behavior-supported cells, with the zero-cell convention used for .
For comparison with the clipped estimator in Definition 16, we also record the unclipped average at the fixed-radius balancing depth. This object is useful for interpreting the bias–variance calculation, while the risk bounds below are stated for the clipped estimator.
Fix and . Let be the minimax squared-error risk in Definition 14, let be the overlap-distance radius in Definition 19, and set There exist constants , with , and an integer , all depending only on and , such that for every and every , and the radius-adaptive partial-history importance-weighted estimator with depth from Definition 20 satisfies Moreover, for every sequence , if is eventually bounded, then there is a constant , depending on the sequence and on , such that eventually
⊢ LeanHolding the radius at a nonunit value gives the fixed-radius slice of Theorem 1. The next theorem states that slice in the fixed-depth notation of Definition 17, together with the signed-depth alternatives from Definition 18 that witness the converse.
Fix . Assume:
(Positive memory exponent.) .
(Positive overlap exponent.) .
(Fixed latent-overlap radius.) .
Let the exponent be There exist constants and an integer such that and, for every integer , where is the minimax squared-error risk in Definition 14. Moreover, the clipped partial-history importance-weighted estimator , with as in Definition 17, satisfies with the stationary target-policy mean reward in Definition 12.
For each integer , there is an integer such that the lower bound is witnessed by the two signed-depth alternatives , , generated by Definition 18 with one observed state and latent states: for every measurable observable-data estimator , under the usual squared-error integrability conditions for both alternatives,
⊢ LeanThe exponent reflects the balance between geometric forgetting and the variance cost of multiplying observable policy ratios. Since is a positive constant on this slice, the all-radius expression in Theorem 1 reduces to the power up to constants. The estimator achieves that rate, and the two-point construction shows that every observable-data estimator faces the same order of risk over the class.
The unit-overlap boundary is the slice of the same surface. At this endpoint, the normalized radius is , and the risk expression in Theorem 1 is on the scale.
Let and . At unit latent overlap, , the following assertions hold.
For every trajectory length , every finite observed and hidden state cardinalities, and every experiment in the model class of Definition 10, the target stationary law and the behavior stationary law of the corresponding policy-induced joint-state transition kernels coincide:
Moreover, there is a constant , depending only on and , such that for every integer , the worst-case squared risk of the immediate observable estimator satisfies Finally, for the minimax squared-error risk of Definition 14, there are constants with and an integer such that, for every , The hidden-state exponent at these positive scales satisfies
⊢ LeanThe boundary result identifies the equality case as an immediate-weighting regime. The coincidence gives the stationary invariance needed for depth zero, and the matching minimax bounds place the endpoint on the scale.
The remaining regime reads Theorem 1 along a deterministic path approaching the unit-overlap boundary. Write the nonnegative excess as ; then traces a local path into the nonunit region.
The following conditions hold:
(Fixed smoothness.) The constants satisfy and .
(Shrinking radius.) The deterministic radius excess sequence satisfies for every .
(Local limit.) The sequence satisfies .
(Rate exponent.) Let .
(Local rate and depth.) Define and
Then there exist constants with , depending only on and , such that, for all sufficiently large , where is the minimax squared-error risk in Definition 14. Moreover, the observable partial-history importance-weighted estimator of Definition 16 satisfies for all sufficiently large . Finally, if there is a constant such that for all sufficiently large , then there is a constant such that for all sufficiently large .
⊢ LeanThe local rate decomposes into the parametric term and a hidden-state term scaled by the shrinking excess. When remains bounded, the immediate estimator already attains the parametric order. When the product grows, the depth in the theorem increases logarithmically with the effective radius , matching the same risk expression.
A finite insulin-policy adaptation
The minimax characterization in Theorem 1 is stated for an abstract finite partially observed model class. This section records a finite insulin-policy construction that keeps the glucose-threshold structure familiar from mobile-health and diabetes treatment applications while putting every component on a finite alphabet (Maahs et al., 2012; Liao et al., 2021; Luckett et al., 2020). The construction, denoted , uses refreshed diet and activity coordinates to create latent heterogeneity, uniform mixing, and stationary overlap in a single controlled Markov model.
The finite insulin-grid state is so the state space has states. The diet coordinates form and all other coordinates form . The behavior and target policies are With probability , the next full state is drawn uniformly. With the remaining probability, draw draw draw , and set The memory coordinates then shift to The reward is This construction defines .
⊢ LeanTwo conventions in the displayed transition fix the numbers the certificate uses, so we state them explicitly. First, is the centered Gaussian of variance , that is standard deviation . Second, the composite clip-and-round operator is the nearest-grid map with cutpoints at the midpoints of the glucose grid, the two extreme cells absorbing everything beyond them; each cell is the interval closed on the left, so a value landing exactly on a midpoint is assigned to the higher grid point, a tie that has probability zero under the continuous noise. Writing for the deterministic part of the displayed expression and for the standard normal distribution function, the next-glucose weights are therefore and, for ,
The observed coordinate contains the glucose level, lagged activity, and lagged treatment, while the diet coordinates are latent. The target rule is a deterministic threshold policy on the observed glucose coordinate, and the behavior rule randomizes treatment with fixed probability. The refresh step supplies a direct minorization component, so the example aligns with the contraction and overlap conditions used by Theorem 1 within a finite-state adaptation of the insulin-policy setting.
The next two facts record the mixing and stationary-law ingredients behind the certified contraction and stationary-overlap constants. They apply to any policy rule on the observed grid states, including the behavior and target policies in Definition 24.
Let and let assign to each of the observed states of the refreshed finite insulin adaptation of Definition 24 a probability law on the action set . Write for the -point joint state space, for the observed coordinate of a grid state , for the transition kernel of , and for the induced transition matrix. Then, for all probability vectors on , equivalently, in the norm ,
⊢ LeanThe proof is deferred to Section B.
Let and let assign to each observed state of the refreshed finite insulin adaptation of Definition 24 a probability law on . Suppose that the induced transition matrix on the joint state space satisfies for all probability vectors on . Then the stationary law selected for is a probability vector and satisfies the stationarity equations
⊢ LeanThe proof is deferred to Section B.
Lemma 1 gives the one-half total-variation contraction supplied by the refresh component. Lemma 2 then identifies the selected stationary distribution as a genuine invariant law for each policy-induced transition matrix, providing the stationary objects needed for the target value and the bias diagnostic.
The reward in Definition 24 is a function of the current grid state. The following regressions make this explicit under target-policy averaging and under one-step target-to-behavior reweighting.
Let , let be the refreshed finite insulin adaptation of Definition 24 with transition kernel and target policy , and for a grid state write for its observed coordinate and for the insulin reward vector. Writing for the numeric reward coordinate of a reward/state pair drawn by , we have, for every ,
⊢ LeanThe proof is deferred to Section B.
Let , let be the refreshed finite insulin adaptation of Definition 24 with transition kernel , behavior policy and target policy , and for a grid state write . Let and let be the insulin reward vector Writing for the numeric reward coordinate of a reward/state pair drawn by , we have, for every ,
⊢ LeanThe proof is deferred to Section B.
Together, Lemmas 3 and 4 show that the immediate reward regression is the same state reward under target averaging and under one-step reweighting. The remaining discrepancy in current-action weighting therefore comes from the stationary law used to average this regression: behavior-stationary weighting versus the target-stationary value in Definition 12.
The finite diagnostic uses certified interval arithmetic for stationary reward averages. The next statements first give the general enclosure principle and then instantiate it for the behavior and target stationary reward averages on the insulin grid.
Let be a finite set, and let be a stochastic matrix on whose entries satisfy for all . The intervals are rational intervals with , combined by exact interval arithmetic: and is the interval whose endpoints are the smallest and largest of the four endpoint products. For an interval , write and for its endpoints and Suppose the following conditions hold.
(Initial row and interval rows.) There are an exact rational probability vector on , a number , and interval rows on such that for every .
(Chunked recurrence.) For every and every output state , there are pairwise disjoint blocks covering , each with at most a fixed positive number of states; block intervals ; and partial-sum intervals such that for each , for , and
(Reward bounds.) There are a reward vector , rational intervals , and a rational bound such that for every .
(Contraction and stationarity.) There are a rational coefficient and a probability vector such that for all probability vectors on , and
Define and Then
⊢ LeanThe proof is deferred to Section B.
Lemma 5 converts a finite number of interval matrix calculations into a bound on the stationary reward average. In the insulin grid, it supplies the certificate used to compare the behavior-stationary immediate-weighted average with the target-stationary value.
For , let and be the terminal and successor rational interval rows on the -state grid of Definition 24 recorded by the zero-step chunked iterate certificate, in blocks of at most states, for the rational interval enclosure of the behavior-policy case or target-policy case . Let , where is the grid reward from Definition 24. Using the exact rational interval arithmetic of Lemma 5, define Equivalently, Then the endpoints of this rational interval satisfy
⊢ LeanThe proof is deferred to Section B.
Lemma 6 is the finite arithmetic step behind the diagnostic. It encloses the behavior- and target-stationary reward averages in rational intervals and shows that their interval difference lies strictly below zero.
The first diagnostic isolates the stationary bias of current-action weighting in this finite construction. It compares the one-step reweighted reward under the behavior stationary law with the stationary target value.
For every natural horizon , let be the stationary immediate-weighting bias for the refreshed finite insulin adaptation of Definition 24, where is the stationary law of the behavior-policy transition kernel on the insulin grid, is the refreshed insulin-grid transition kernel, and are its behavior and target policies, is the one-step target-to-behavior policy ratio, and is the target value in Definition 12. Then
⊢ LeanThe proof is deferred to Section B.
The interval in Lemma 7 is a finite-state certificate for the immediate-weighting bias diagnostic . The theorem below records the model-class properties of the same construction and restates the diagnostic in the notation of the insulin-grid transition law.
For every horizon , the refreshed insulin experiment of Definition 24 satisfies:
(State count.) Its observed–latent joint-state alphabet has cardinality .
(Uniform contraction.) It satisfies Assumption 6 with contraction parameter .
(Policy overlap.) It satisfies Assumption 5 with one-step overlap constant .
(Stationary overlap.) It satisfies Assumption 7 with stationary-overlap constant .
Define the immediate-weighting bias diagnostic by where are the transition kernel, behavior policy, and target policy of , is the observed coordinate of , and is the stationary law induced by . This diagnostic is interval-certified as strictly negative:
⊢ LeanTheorem 5 certifies five things about the grid, and exactly these five. Its joint observed–latent alphabet has states. It contracts in total variation at rate , so it meets Assumption 6. Its policy pair has one-step overlap constant , so it meets Assumption 5. Its two stationary joint-state laws satisfy the latent stationary overlap bound of Assumption 7 at radius . And its immediate-weighting diagnostic lies in the interval , certified by interval arithmetic.
The last of these is the substantive one, and it is what the construction was built to isolate. Immediate weighting reweights the observed reward by the current-action ratio alone, so it averages the observed-glucose reward under behavior-stationary occupancy; the target value averages the same reward under target-stationary occupancy. The certified interval quantifies the gap between these two occupancies in a model whose latent stationary overlap is finite and whose mixing is fast. Bounded latent occupancy and fast forgetting are therefore compatible with a stationary bias for current-action weighting, and the source of that bias is occupancy rather than slow mixing or a divergent action ratio. This is the mechanism the depth parameter of Definition 16 exists to control.
Discussion and limitations
Theorems 2, 4, and 1 describe a risk surface indexed by the stationary latent-overlap radius. For each fixed , the minimax risk has rate , while at it has the parametric scale. The radius-sensitive statements refine this picture by expressing how the risk changes as approaches one. In that local regime, the normalized radius from Definition 19 is the relevant coordinate: it measures how far the target stationary occupancy may move from the behavior stationary occupancy while retaining the finite-state POMDP restrictions of Definition 10.
The elbow has a direct statistical interpretation. Below this scale, the local overlap distance is small enough that the parametric floor governs the risk, as recorded by Theorem 3 and the small-radius branch of Theorem 4. Above this scale, the hidden-state occupancy discrepancy becomes visible in the observable-data experiment through longer histories, and the larger-radius branch of Theorem 4 gives the corresponding degradation. The radius-calibrated history depth in Definition 20 operationalizes this tradeoff: it selects depth zero in the parametric neighborhood and increases the observable history length when the overlap radius makes latent imbalance statistically consequential.
The role of latent stationary overlap is therefore distinct from ordinary one-step action overlap. The one-step condition controls the variance of products of observable policy ratios, while the stationary density-ratio condition controls how target and behavior policies populate hidden states. Theorem 1 is the master rate statement: Theorems 2, 3, and 4 are the fixed-overlap, unit-overlap, and local shrinking-overlap slices. The proof balances the PHIW bias scale , coming from contraction after conditioning on a length- observed history, against the variance scale , coming from products of one-step policy ratios. The signed-depth lower bound chooses a terminal depth so that the observable KL is controlled by a constant times , while the target value gap is calibrated to the upper-bound bias scale. The radius-calibrated depth in Definitions 20 and 16 attains the displayed frontier when calibrated with , and a known shrinking law for gives the corresponding changing depth.
The finite insulin demonstration in Theorem 5 is a controlled finite-state policy example. The construction supplies a refreshed Markov model with finite contraction, one-step policy overlap, and the sufficient stationary density-ratio bound , hence . Its immediate-weighting diagnostic remains separated from zero under the displayed interval arithmetic because current-action weighting averages the observed-glucose reward under behavior-stationary occupancy rather than target-stationary occupancy. The example therefore exhibits, in a stylized treatment-policy model, the occupancy mechanism that the depth parameter of the partial-history estimator is designed to control. Its certified content is the four structural constants and the bias interval; establishing the full model-class predicate for this grid — in particular fixing an initial law and verifying the stationary-start, Markov-kernel, sequential-ignorability, and bounded-reward conditions at and — is a separate verification we do not carry out here, so the frontier theorem is not being applied to this grid.
Limitations.
The signed-depth converse in Definition 18 uses a latent alphabet whose terminal depth grows with , so the lower-bound construction is calibrated to a sequence of finite experiments. Confidence intervals and policy-learning regret are separate research directions built on top of the value-estimation problem studied here. The finite insulin adaptation in Definition 24 is a controlled finite-state demonstration designed to instantiate the paper’s overlap and contraction constants within an interpretable treatment-policy example. Three limits of that demonstration are worth stating. Its certificate covers the four structural constants and the bias interval, not the full model-class predicate, so no rate guarantee of this paper is asserted for it. It is not a finite-sample comparison of estimators, and it does not show that the depth calibrated from the sufficient bound is the practically preferable depth on this grid. And a data-driven choice of depth for an unknown radius remains open here as elsewhere in the paper.
Appendices
Proofs and auxiliary constructions
This appendix collects the auxiliary statements used to support the main rate calculations. The first group records the observable filtration and the finite-prefix quantities that enter the signed-depth likelihood comparison.
Define and where the empty product is equal to one.
⊢ LeanThe quantities in Definition 25 separate two sources of information in the observed history. The posterior captures the chance that the hidden depth has reached its terminal value, while the factor records the visible action pattern needed for that event under the rare-action construction. This separation is the device that makes the observed likelihood calculation finite-prefix and observable.
The upper-bound argument begins with local second-moment and covariance controls for the PHIW scores. These statements isolate, respectively, the contribution from a single weighted window and the dependence between overlapping windows.
Let , let , let , and let be a raw POMDP experiment on finite observed and hidden alphabets. Suppose that:
(Sequential ignorability.) satisfies Assumption 2.
(Policy overlap.) satisfies Assumption 5 with constant .
(Overlap level.) .
(Kernel law.) satisfies Assumption 1.
(Bounded rewards.) satisfies Assumption 4.
(Window position.) and .
For the depth- score of Definition 16, under the observed trajectory law of ,
⊢ LeanWrite an observed trajectory as and set For every , define the observable zero-cell policy ratio Then Definition 16 gives Since , the window is the contiguous block
The window product has observed-law mean one: Put . Since , this is an integer in , and the window from step 1 is exactly the interval . First establish the corresponding identity under the full trajectory law. For a full trajectory , write so the full-law block identity to prove is The argument uses the theorem hypothesis that the target carrier is a policy vector, so for every observed state , together with the domination inequality supplied by Assumption 5. Under Assumption 2, conditional on the past and current joint state at epoch , the action has behavior probabilities . If , then , hence ; therefore, for every observed state , Thus, for every , conditioning on the history and current joint state up to epoch gives Peeling the factors one epoch at a time, with Assumption 1 identifying the next history law after each peeled action, gives where the product over is the empty product. Iterating from down to leaves the empty product, whose expectation is , and hence It remains to pass this identity to observed trajectories. Let be the observation map so that the observed law is the pushforward of the full law by . By definition of and the interval identity , for every full trajectory , The usual integration formula for a pushforward measure therefore gives
For every observed trajectory and every , The lower bound uses the behavior-policy vector property from Assumption 2 and the theorem hypothesis that the target carrier is a policy vector. The upper bound is immediate in a zero cell and otherwise follows from the domination inequality in Assumption 5 by dividing by the positive number . Hence The map injects into , so Since ,
By Assumption 4, the full trajectory law has almost surely; pushing forward to the observed law gives Together with the preceding product bound, this also gives so the square and the comparison function are integrable under the observed law.
For -almost every observed trajectory, where the last inequality uses . Integrating and using the normalization from step 2,
Let be a raw POMDP experiment of length on finite observed and hidden alphabets. Let , let , and let . Assume:
(Model conditions.) satisfies Assumption 2, Assumption 5 with overlap constant , Assumption 1, and Assumption 4.
(Overlap constant.) .
(Initial window.) .
(Lag identity.) .
(Overlapping lag.) .
Then the depth- scores of Definition 16, evaluated under the observed trajectory law of , satisfy
⊢ LeanLet be the observed-trajectory law of . Since it is the pushforward of the probability law of the full raw experiment under the observation map, is a probability measure. Also, from , , and , we have
We use the following covariance reduction. If is a probability law, , and a number satisfies then Indeed, by expanding the covariance and applying the triangle inequality. The two -bounds give because both factors are nonnegative. Hence
Apply this reduction with for observed trajectories . The two scores are square-integrable under . To see this, write the one-step ratio in Definition 16 as for every observed trajectory . The policy-vector parts of Assumptions 2 and 5 give . If , Assumption 5 gives If , then , hence , and the zero-cell convention sets the ratio equal to zero. Thus, for every observed trajectory , Also almost surely by Assumption 4. Hence, for every time , outside the null set where the reward bound fails. The window is a subset of the integer interval , so . The map is measurable: the observed trajectory space carries real rewards and is not finite, so measurability is not a counting argument, but each coordinate map and is measurable, the ratio is a measurable function of the first of these by the two-case definition above, and is a finite product of these measurable maps. Since is a probability law, the almost-sure bound above implies , and in particular it applies to and .
Moreover, Assumption 1 and the initial-window conditions and give For a single time with , this last estimate is obtained as follows. For integers with , define with the empty product equal to one when the lower endpoint is above the upper endpoint. This chronological block has unit expectation: To prove this, pull the integral back from to the full trajectory law, so that , and are full-trajectory coordinates. For , let and where the product is one for . The factor is -measurable. By Assumption 2, conditionally on the action has law . The zero-cell convention and the implication computed above give and hence The remaining coordinates after the action at time are integrated with the probability kernel in Assumption 1; integrating first over the fibre containing preserves the expectation of any function of . Applying the displayed conditional identity successively for removes the factors one at a time and leaves the empty product, whose expectation is one. Taking and gives Now and almost surely by Assumption 4. Therefore
It remains to bound the cross moment. For the two overlapping windows define Since and are elements of , their integer values are valid finite-time endpoints. The assumptions , , and give and identify the union in the formal finite index set as The same endpoint inequalities give the cardinality bound for the overlap, Using the same zero-cell ratio as above, the policy-vector and overlap parts of Assumptions 2 and 5, together with , give Therefore the second inequality by the cardinality bound together with the monotonicity of for ; the intersection need not have exactly indices, so this step is an inequality and not an identity.
Since and almost surely, The product over two windows decomposes into the product over the union times the product over the overlap: Combining this with the previous display gives the pointwise almost-sure bound The finite-window identity from the preceding step rewrites the union product in exactly chronological form: The interval-normalization argument proved above applies with lower endpoint and upper endpoint . The lower endpoint is nonnegative because , the upper endpoint is in because is a time index, and follows from . Thus Hence
Finally, because . The covariance reduction from Step 2, applied with , yields which is the claimed bound.
The analysis of separated windows uses both the observed law and the full trajectory law. The next definitions fix those probability spaces and the observation map connecting them.
For a raw POMDP experiment , let denote the probability law of the observed trajectory induced by the behavior-policy experiment . Equivalently, under , , , and generates from , with only retained.
For a raw POMDP experiment with joint state space , let denote the full behavior-policy trajectory law on Under , the initial state satisfies , each action is drawn as , and the kernel generates from , for .
Let be the observation map from a full trajectory to its observed trajectory, where . The observed behavior-policy law is the pushforward
Together, Definitions 26 and 27 let the covariance calculation move between observable scores and latent-state transition operators. The observed law is the sampling law for the estimator, while the full law retains the latent state needed to apply contraction after a gap.
Let be a raw POMDP experiment with trajectory length , and let , the policy-overlap constant, be real. Let denote the depth- partial-history weighted reward score from Definition 16. Suppose:
(Model conditions.) satisfies Assumptions 2, 5, 1, and 4 with overlap constant .
(Overlap scale.) .
(Mature earlier window.) The history depth satisfies .
(Separated future window.) For an epoch and a gap , the epoch lies in the observed horizon, equivalently .
For , let be the policy-induced transition operator obtained from the reward-marginalized kernel under policy , let and put Then, with the observed law induced by the observed trajectory map ,
⊢ LeanSet The horizon condition gives , and makes the earlier score mature. For a full trajectory , write the joint state at epoch as and write for the state-and-epoch history through epoch . For an observed trajectory , write its epoch- coordinates as Define the totalized one-step ratio, for an observed state and action , by For , put Thus the superscript means that the upper endpoint is excluded: the largest ratio in a nonempty block is . In particular, stops at , while the terminal factor is written separately below. The target-policy one-step reward regression is and the separated future function is . In this proof the policy-induced transition operator is the reward-marginalized kernel for every bounded .
Since is the pushforward of under , The measurability needed for this change of variables follows from the observable definition of in Definition 16.
Define lifted carriers on state histories by and, for , let be the lift of that forgets the last appended action, reward, and next state: The base carrier is well defined because uses only rewards, actions, and observed states through epoch , while ; hence every lifted carrier satisfies The depth- window in is inclusive. Since and , the half-open block contains the factors with , and the endpoint contributes the separate terminal factor: Therefore The totalized ratios obey : on a positive behavior cell this is the domination inequality divided by , and on a zero cell it follows from the definition of . Hence finite state spaces, Assumption 4, and the finite product bound make all displayed integrands integrable. Empty behavior or target blocks are interpreted as products equal to one; in particular, when the block is empty and the endpoint factor is still present.
The terminal reward peel is Indeed, condition first on the history and action at . Assumption 1 supplies the conditional law of as . Applying this conditional-law identity to the measurable test function , with integrability supplied by Assumption 4, and then marginalizing over the successor state gives By Assumption 2, the action is drawn from . For every observed state and action , the totalized ratio satisfies Indeed, if , this is the definition of . If , then the domination inequality in Assumption 5 gives , while the target policy has nonnegative coordinates, so , matching the zero-cell value. Applying this identity at and summing over gives by the displayed definition of .
For every , put , so . For every bounded , the target-ratio peeling step is The first target peel has and lowers the carrier from to ; the last has and lowers to . Iterating the displayed identity backward over these values of therefore removes exactly the target-window ratios. Indeed, is the lift of , and Conditioning on the history and action at , Assumption 1 gives the transition integral over , and the reward coordinate is integrated out: The ratio identity displayed in the terminal peel justifies the replacement of by , including zero behavior cells, and the last display is exactly . Iterating for , starting from , yields When , this iteration is empty and the displayed identity reduces to the same integrand at .
For every , put . For every bounded , the behavior-gap peeling step is Here is the history obtained from by appending , and is the lift of , so . Conditioning on the history and action at , Assumption 2 supplies the behavior action law, while Assumption 1 supplies the joint law of given . Integrating out the reward coordinate gives, for each history through , which is the reward-marginalized behavior transition appearing in the displayed identity. Iterating over with gives When , this iteration is empty. Since , , and , combining the preceding displays proves
Lemma 10 rewrites a separated score product as an earlier observable score multiplied by a future continuation value evaluated at . This identity supplies the point where the gap between windows becomes a behavior-transition power, followed by a target-policy window.
Let satisfy Definition 10, with behavior policy , target policy , and the policy-overlap condition in Assumption 5. For , write and let and . Let denote the behavior stationary law of .
Assume:
(Policy choice.) The policy is either or .
(Terminal agreement.) The functions satisfy for every joint state with .
(Initial support.) The joint state satisfies .
Then, for every ,
⊢ LeanFix a policy . The argument first records the support-closure fact used at the induction step. For joint states and an action , write Then
Suppose and . We claim that First show that . If , this is immediate. If , suppose instead that . By Definition 10, the behavior policy is a probability vector and is a Markov kernel, so every summand in is nonnegative. Hence, for each action , For such an , if then If , the preceding zero product gives ; by Assumption 5, and the target policy is nonnegative by Definition 10, so . Thus again Summing over gives contradicting in the case . Therefore .
By Definition 10, is stationary for , so All terms in this sum are nonnegative. Therefore which proves the claim.
We now prove the asserted identity by induction on . For , because and on the behavior support.
Assume the statement holds for : for every joint state with , Let satisfy . Using the recursive definition of the iterated operator, Fix . If , then the two -summands are both zero. If , the row nonnegativity of gives . The support-closure claim from the previous step then gives , so the induction hypothesis yields Multiplying by , the two -summands agree. Since this holds for every , the sums agree: The induction proves the result for every .
The support congruence in Lemma 11 keeps transition-operator calculations on the behavior support. Because the overlap condition makes target transitions compatible with behavior support, functions that agree on the behavior-stationary support continue to agree after iterating either policy operator from a supported state.
Let , let , and let be a raw POMDP experiment as in Definition 10, with the uniform contraction restriction of Assumption 6. Write the contraction factor as . Let satisfy Then the depth- scores of Definition 16, evaluated under the observed trajectory law of , satisfy
⊢ LeanPut Then by , and Let denote the full trajectory law and its pushforward under the observation map . For , write and define
Since and in Definition 10, . The maturity and separation hypotheses give , so Lemma 10 applies and yields For the companion mean identity, let and , so . Since , the depth- window in is exactly . We claim To prove this identity, condition first on . Write The zero-cell convention and Assumption 5 imply whenever . Hence, for each pre-terminal epoch , while at the terminal epoch Thus the target-weighted terminal window has conditional mean , with the terminal display alone applying when . Conditioning backward over the gap epochs , no score ratio appears, and Assumptions 2 and 1 give the behavior propagation ; when , this is the identity operator. Bounded rewards and the finite state and action spaces justify the displayed integrals. Therefore the conditional mean given is , proving the claim. Also, by definition of , Thus the covariance equals the centered full-law product The covariance identity is legitimate because Indeed, write the window product in as The product contains at most factors. The zero-cell convention and Assumption 5 make each factor nonnegative and at most on observed behavior-support cells, while Assumption 4 gives almost surely. Hence almost surely under , and the two scores are square-integrable.
The first score has unit norm: Indeed, with the zero-cell convention and Assumption 5 give by peeling the ratios from down to : conditionally at each peeled epoch, By Assumption 4, almost surely, which proves the displayed bound after passing between the observed law and the full law by the observation map.
Let be the stationary law of , and define the support-truncated regression on all of by For every with , To see this, fix such an . For an action with , consider epoch . The history before epoch is the unique empty history, so the full history-action singleton with current state and action has probability by Assumption 3 and the epoch-zero action factorization in Assumption 2. Under the history-next-pair factorization supplied by Assumption 1, Assumption 4 gives, for almost every history-action pair, a kernel row whose reward coordinate is almost surely in . Since this epoch-zero singleton has positive mass, that almost-sure row statement holds at , and hence If , then Assumption 5 forces . Therefore For , the definition gives . Consequently
Since , satisfies . For a function , write For each , the rows of are probability vectors. Nonnegativity follows from the nonnegativity of the policy probabilities and of the kernel rows. For the row sums, Here Assumption 1 supplies the kernel row mass, Assumption 2 supplies the behavior probability vector when , and the target-carrier part of Assumption 5 supplies the target probability vector when .
We first convert the total-variation contraction in Assumption 6 into the oscillation contraction used below. Let and be probability vectors on the finite state space, and let . Put Then for every , and, since and have equal total mass, The layer-cake identity on the finite state space gives and therefore Now fix states . Applying Assumption 6 to the point masses and gives Combining this with the preceding dual bound yields the one-step estimate Taking the supremum over and iterating gives the finite-state oscillation form: whenever for some , Applying this first with , , and , and then using , gives Applying it next with to gives
Define Because on , Lemma 11 applied with policy gives, for every behavior-supported state , This pointwise behavior-supported equality is the terminal-agreement hypothesis needed for the second application of Lemma 11, now with policy . Thus, for every initial state with , We next record the behavior-stationarity identity used at time . For every bounded and every epoch , This follows by induction on . At the initial epoch the identity is Assumption 3. For the induction step, Assumption 2 and Assumption 1 give the behavior peeling identity Using the induction hypothesis with the bounded function , the right-hand side becomes where the last equality is the stationarity equation for the behavior stationary law , part of the model-class data in Definition 10. Taking the zero-support indicator at gives so for -almost every . Hence
Put Using the almost-sure replacement from the previous step, both and hold; the integrability needed for these replacements follows from the finite state space and the score bound above. The centered representation from the first step therefore gives The oscillation bound in the fourth step implies, for every , because is an average of values of . Therefore, using the unit bound from the second step,
Together, Lemmas 8, 9, 10, 11, and 12 decompose the sampling fluctuation into a weighted-window term and a geometric dependence tail. The overlap condition controls the inflation inside the window, while contraction controls pairs whose windows are separated by at least one time step.
The next statement sums those covariance bounds over future times. It is the finite-sample accounting step that turns the local bounds into a variance inequality for the average score.
Let , , , and let be as in Definition 10. Write , , and . Then, for every , the depth- scores of Definition 16 satisfy, under the observed trajectory law of ,
⊢ LeanFix . By Definition 10, and , so The same membership supplies the model restrictions used by the covariance estimates below.
Define the future index set and its two lag classes by and For each , the lag is a positive integer, hence exactly one of and holds. Thus Writing we get
For , the lag identity is and . Since and , Lemma 9 gives Therefore The map is injective and takes its values in . Since , For , Combining the last three displays yields
For , the lag identity is and . Hence Lemma 12 gives Thus The map is injective and takes its values in . Since , Because , Consequently,
Adding the two bounds over the disjoint decomposition of gives
Lemma 13 supplies the two terms that reappear in the assembled variance bound. The geometric sum over overlapping lags gives the policy-overlap contribution, and the contraction tail gives the dependence contribution.
Let and be nonnegative integers, let , and let be a raw POMDP experiment. Write for the policy-overlap factor and for the contraction factor. Suppose that
(Horizon.) .
(Model class.) , with as in Definition 10.
(History depth.) .
For the depth- scores of Definition 16, form the unclipped partial-history average Under the observed trajectory law obtained by projecting the full trajectory law of onto observed histories, its variance satisfies
⊢ LeanLet denote the observed trajectory law and let denote expectation under this law. Let denote the full trajectory law carried by , and let denote expectation under that law. Put For , write Since , Definition 10 gives , , hence Also implies In particular . The raw average is
For each , Lemma 8 applies with the sequential-ignorability, policy-overlap, kernel-law, and bounded-reward clauses supplied by Definition 10, together with and . It gives Using the standard identity we obtain
Fix . The aggregate covariance estimate in Lemma 13 applies to the same model class and the same future index set . Therefore
Now expand the variance of the finite sum: Bounding each covariance by its absolute value and then using covariance symmetry gives Using the diagonal and future-sum bounds,
Finally, Since , , and , Substituting the definition of gives
The two terms in Lemma 14 have distinct origins: multiplying observable policy ratios over a window produces the term, while serial dependence contributes the contraction term. The bound is stated for the unclipped average because the displayed variance is the component used before the clipped estimator’s squared-error risk is assembled.
The next statement supplies the fixed-radius upper bound for the partial-history importance-weighted estimator. It gives both the rate bound for the balanced depth and the corresponding variance calculation for the unclipped average.
Fix constants satisfying
(Mixing scale.) .
(Policy-overlap scale.) .
(Latent-overlap radius.) .
Let be the balanced history depth in Definition 17, and set Then there exist constants and such that, for every , the estimator-specific worst-case squared risk from Definition 13 of the observable PHIW estimator from Definition 16 satisfies Moreover, for every , every finite alphabet sizes , and every raw POMDP experiment , the variance under the observed trajectory law induced by of the unclipped partial-history average satisfies
⊢ LeanDefine The assumptions and give Also implies Let Then and Since , there is an integer such that, for every , For such , Definition 17 gives and hence Therefore and With we have, for every ,
Set and define Since and , the constants are positive, and so .
Fix . Then , , and . For every , Lemma 34, applied with , gives Using , , and , The balance bound from the first step and the inequality , valid because and , yield There are two cases. If the indexed class is nonempty, the pointwise bound over every member of the class, together with the estimator-specific risk definition in Definition 13, gives If the indexed class is empty, the real-valued worst-case-risk functional used for has empty supremum equal to . Since , , and , the bound also holds in this case.
It remains to record the variance assertion. For every , the definition of gives , and the minimum in Definition 17 gives . Apply Lemma 14 with , , and . Its conclusion is exactly uniformly over the finite alphabet sizes and every , as required.
The bound in Lemma 15 identifies the observable estimator used for the upper side of the fixed- comparison. The variance display isolates the cost of multiplying observable policy ratios over a window of length and the serial-dependence contribution controlled by contraction. Balancing those terms gives the exponent under fixed overlap and fixed latent-stationary domination.
For the lower side, the appendix uses the signed-depth alternatives introduced in Definition 18. The next two statements record the contraction and stationary-law calculations that make the construction a member of the same finite-POMDP class.
Let , , , , and , and write . Let be the transition kernel induced on the joint state space of the signed-depth alternative of Definition 18 with terminal depth and index by the constant policy that plays action one with probability . Then, for all probability vectors on ,
⊢ LeanLet For and in , define the refresh law and define the residual kernel Expanding the transition rule in Definition 18 under the constant policy , , gives, for every , Indeed, at terminal depth both sides equal ; at depths , the reset part contributes , while the advance-to- part contributes .
The residual kernel is stochastic. Since , one has , so the two reset weights are nonnegative and If , summing over therefore gives . If , the only nonzero entries have , and their two sign weights sum to with nonnegativity following from . Thus for every .
For any signed vector on , stochasticity of gives the total-variation nonexpansiveness bound
Now take probability vectors . Because , the common refresh term cancels: Since , the preceding display and nonexpansiveness yield This is the claimed contraction.
∎Let , , , , and . Write , , and . Let be the kernel induced on by the constant policy for the signed-depth alternative of Definition 18 with terminal depth and index , and let be the stationary law selected for . Then, for every hidden state ,
⊢ LeanWrite We prove that this probability vector is stationary for , and then use the contraction statement to identify it with the selected stationary law.
Since , the value lies in . Also gives , and gives . Hence every factor is nonnegative. For each fixed depth , and the geometric normalization gives Thus is a probability vector on .
The state kernel induced by the constant policy has the following hidden-state transition probabilities, obtained from the signed-depth dynamics in Definition 18 by averaging the action-one and action-zero advance weights. If , then and If , then The remaining transitions have weight zero.
For a target state of depth , summing the displayed reset probabilities gives Since , Therefore
For a target state of depth , only depth can advance to it, so The sign sum is Using , this yields Thus .
The contraction supplied by Lemma 16 implies uniqueness of stationary probability vectors. Indeed, if and are stationary for , then Because and total variation is nonnegative, the norm is zero, hence on the finite state space. The selected law is stationary because the stationary vector exists, so uniqueness gives .
Evaluating the identity at the hidden state gives as claimed.
Lemma 16 gives the uniform total-variation contraction of the signed-depth transition under any constant action-one probability. Lemma 17 then gives the stationary law explicitly: the depth mass is a truncated geometric sequence, and the sign imbalance is scaled by the action-one probability to the current depth.
For every , , , , , and , let be the signed-depth experiment of Definition 18. Put where is the mixing factor and . Then:
(Membership.) in the sense of Definition 10.
(Behavior stationary law.) For every hidden state , with decoded depth and sign ,
(Constant action-one stationary laws.) For every , if denotes the stationary law induced by the constant policy that chooses action one with probability , then for every hidden state ,
(Target-over-behavior domination.) For every joint state ,
The stationary target value of Definition 12 is
⊢ LeanWrite , , and . Since and , we have and . Also gives , so For a hidden state , abbreviate and .
For , let be the constant policy with , and let be the induced transition kernel on the joint state space By Lemma 16, for all probability laws and on , The behavior policy in Definition 18 is the case , while the target policy is the case . The displayed inequality therefore gives the contraction required in Assumption 6 for both policies.
For , Lemma 17 identifies the selected stationary law of the constant-policy kernel : for every hidden state , Taking gives the behavior stationary law because the behavior policy in Definition 18 chooses action one with probability , and . Thus Taking gives the target stationary law
We now verify the model-class predicates in Definition 10. The embedded experiment is generated by the initial law, the behavior policy, and the time-homogeneous reward-transition kernel in Definition 18. Therefore, for each epoch , which is Assumption 1. The behavior action law is and these two masses are nonnegative because . Since the observed state is the singleton , the chronological law factors at every epoch as which is Assumption 2. The finite reward symbols are decoded as and , so giving Assumption 4. The policy-overlap inequality is so Assumption 5 holds with factor . The initial law is where the last equality is the stationary-law formula from the case . Thus Assumption 3 holds.
For the latent stationary overlap, first identify the two selected stationary laws and the initial law on the same cells. The target and behavior formulas from Step 2 give The displayed initialization in Definition 18 gives where the last equality is the behavior stationary-law formula from Step 2. The depth mass is nonnegative: since we have Here . Since , we have . For , where the first inequality uses , and the last uses and nonnegativity of the behavior cell. For , because and , so the multiplier is nonnegative, while . Multiplying this last inequality by that nonnegative multiplier gives the displayed inequality, and its right-hand side is by the formula for . Every joint state has the unique observed coordinate , so these two sign cases give Moreover because either , giving , or , which is possible only when . Since , so Assumption 7 holds.
It remains to collect the remaining requirements of Definition 10. The horizon requirement is the hypothesis of the present lemma. The observed alphabet of is the singleton , the latent alphabet is the -element set , and the action alphabet is ; all three are finite, and the action alphabet has exactly two elements, as the class requires. The parameter requirements and are hypotheses of the lemma, and they fix and as the overlap and contraction constants used throughout. Together with the seven assumptions verified above – the Markov kernel, sequential ignorability, stationary start, bounded rewards, policy overlap, uniform contraction, and latent stationary overlap – this is every defining property of Definition 10, and therefore .
It remains to compute the target value. The two-point alternative label enters the reward law only through its numeric sign: the positive label contributes , the negative label contributes , and in the statement this multiplier is denoted by . Under the target policy the action is always one. From the reward weights in Definition 18, Using Definition 12 and the already computed target stationary law, Only depth contributes, and summing over the two signs gives This is the claimed value formula.
The stationary laws in Lemma 18 show how the lower construction keeps the target-over-behavior ratio bounded while concentrating the signal at terminal depth. The terminal value is proportional to the stationary depth mass , so increasing makes the two alternatives harder to separate from observed data while preserving the fixed latent-overlap radius.
The likelihood calculation rests on finite-prefix posterior bounds, a rare-action second moment, a bootstrap factor that makes the posterior control uniform in the terminal depth, signed-mass identities at terminal depth, exact conditional reward means, and finite-word KL inequalities. The next group develops these ingredients in the order in which the observed-path comparison uses them.
Let , let , let , and let be an observed prefix for the signed-depth family of Definition 18. Write and, for , write for the first observed pairs. For any observed prefix , define the raw terminal-depth posterior by Assume:
(Nonnegative amplitude.) .
(Nonnegative base.) .
(Window length.) .
(Strict-prefix posterior bound.) For every ,
Define the final -prefix lower window by Then
⊢ LeanWe prove the slightly stronger assertion with the prefix length left variable. For a prefix of length , write with the convention that an empty product is .
1. If , then the condition selects no indices, so This proves the base case.
2. Assume the claim has been proved for . Consider a prefix of length with . If , this inequality is impossible, so there is no case to prove. Thus write . Let be the length- prefix obtained by dropping the last observed pair: For every , the first observations of are exactly the first observations of : Hence the strict-prefix posterior hypothesis for gives Since , we have , so the induction hypothesis applied to yields
3. The final-window product for length splits into the length- final-window product for and the last factor: Indeed, the threshold satisfies so the eligible indices before the last one are precisely the eligible indices in the length- window, and the remaining eligible index is .
4. The posterior hypothesis at the last strict prefix gives . Because , multiplying preserves the inequality, and subtracting from reverses the side on which the smaller product appears: Together with the assumed nonnegativity , we have Combining this nonnegativity, the induction bound, and the last-factor bound gives This completes the induction, and hence the desired bound for the original prefix of length .
∎Let , let , and let . Put . In the finite signed-depth model of Definition 18, for an observed prefix , write For , define the lower posterior window where is the prefix of containing its first action–reward pairs.
Assume:
(Positive scale.) .
(Positive tilt.) .
(Overlap level.) .
(Mature window.) .
Then the terminal-depth posterior satisfies
⊢ LeanSet , and introduce and Here the products are over the final action–reward pairs; if , they are empty products and equal one.
1. The terminal-depth numerator after the full prefix factors as Indeed, summing over the two terminal signs gives the unsigned mass at depth . Over a mature window of length , the only way to arrive at terminal depth is to start the window at depth and advance once at each of the updates; this contributes , while the action and sign-symmetric factors over the same window contribute exactly . The remaining mass is the unsigned depth-zero mass at the beginning of the window.
2. The total mass over the observed prefix factors through the same window as This follows by iterating the one-step mass recursion across the last observations: each step extracts the behavior/fair-sign factor and leaves the corresponding reward-correction factor. For the window is empty and the identity reduces to .
3. The unsigned depth-zero mass at the window start is bounded by the total mass there: All summands are nonnegative, since each prefix weight is a sum of nonnegative path weights.
4. The lower posterior window is positive and is dominated by the correction window: For each factor in , so . The one-step reward-correction factor at the same prefix is at least and taking products over the final window gives the displayed inequality.
5. The remaining factors have the needed signs: The first inequality is the product of positive behavior weights and factors, while the second is positivity of the finite total mass at the window start. The third follows from Step 4 by . Finally, , so its -th power is nonnegative.
6. Combining the two window factorizations with the monotonicity bounds gives because , , and . Multiplying by the nonnegative factor yields Using the identities from Steps 1 and 2, this is Since and , division gives as claimed.
∎Let , let , let , let , and let be an observed action/reward prefix for the signed-depth construction of Definition 18. Write . For any prefix , define the raw terminal-depth posterior by Define the initial lower posterior window over the available prefixes by where is the prefix of of length .
Assume:
(Positive scale.) .
(Positive overlap exponent.) .
(Nontrivial overlap radius.) .
(Initial window.) .
Then the terminal-depth posterior after observing satisfies
⊢ LeanWrite for , and let denote the empty prefix. For any prefix , set Let be the behavior-policy probability of action from Definition 18, and define For a prefix and a next observation , define the reward-correction factor whenever , with any fixed value when the denominator is zero. Along the observed prefixes below the denominator is positive. Finally set
Since , a hidden path contributing to starts, after summing over the sign coordinate, from depth and then makes successive depth-increment moves. Thus the terminal numerator factors as where is the signed-depth initial depth mass from Definition 18. Since , and therefore
Summing the one-step kernel over the next hidden state leaves total hidden-transition mass one. Hence each extension satisfies Multiplying this identity for , with the empty-product convention when , gives
The empty-prefix mass is one. Indeed, the initial weights sum to Thus
The lower window is positive and is bounded by the correction window: To see this, gives , hence For each observed prefix , the filter weights are nonnegative and their total mass is positive, so Consequently every factor is positive. Moreover, for each , because the doubled reward factor equals off terminal depth and is at least on terminal depth. Multiplying these pointwise inequalities gives the displayed product bound.
The fair-window factor is strictly positive: Each behavior probability is positive under , and each reward-symbol base factor is positive; the empty product is .
The terminal depth mass is bounded by : Indeed, and implies and . After multiplying by the positive denominator, the desired inequality is which is exactly times the nonnegative factor .
Combining the previous bounds, Multiplying by the nonnegative factor , and using the identities above, yields Since and , division gives
Let , let , and let . Suppose that
(Positive temperature.) .
(Positive log-overlap.) .
(Overlap level.) .
For the finite signed-depth model of Definition 18, write For every observed prefix , define the raw terminal-depth posterior by where the sums range over the signed-depth hidden states at depth horizon . Then
⊢ LeanWrite and .
1. Since , one has . Hence For every observed prefix , the raw terminal-depth posterior satisfies Indeed, the numerator is a sum of nonnegative filter masses over terminal-depth states, the denominator is positive, and the terminal-depth mass is bounded by the total mass.
2. First suppose . By Lemma 20, Apply Lemma 19 with and . The hypotheses are exactly those checked above: , , , and for every strict prefix. Thus Since and , division by the larger positive denominator gives
3. It remains to consider . By Lemma 21, Apply Lemma 19 with and . Since , and the prefix posterior bound holds for every strict prefix, this gives Therefore Because and , monotonicity of powers on yields Dividing the nonnegative numerator by these positive denominators gives Combining the displayed inequalities proves the claimed bound in the initial-window case as well.
∎Let be nonnegative integers. Suppose that
(Positive scale.) .
(Positive perturbation.) .
(Overlap level.) .
Write for the contraction parameter and for the signed-depth reward-amplitude constant. For an alternative in the signed-depth family of Definition 18 and any observed prefix , let be the raw terminal-depth posterior: the run-local unnormalized filter mass at latent depth , divided by the total run-local unnormalized filter mass of the prefix. Then
⊢ LeanLet Since , the contraction parameter satisfies . Hence and therefore . Moreover, Thus . Consequently is strictly positive; indeed .
For , define the lower posterior window where denotes the first observed action–reward pairs, with the empty prefix; for this is the empty product and equals . If , Lemma 20 applies under the present hypotheses and gives If , Lemma 21 applies under the present hypotheses and gives These two estimates are exactly the two window cases needed below.
The crude terminal-posterior estimate in Lemma 22 applies to every observed prefix , of any length, under the same assumptions , , and . With the present notation it states This supplies the prefix-uniform bootstrap bound used in the lower-window estimate.
Apply Lemma 19 with Its nonnegative-amplitude hypothesis is , established in the first step, and its nonnegative-base hypothesis is also established there. Its strict-prefix posterior hypothesis holds because the preceding step gives, for every prefix appearing in the window, The indexing in Lemma 19 ranges over the final strict prefixes of , precisely the factors in . Hence, for every ,
If , the terminal-window bound with and the preceding display give If , the terminal-window bound with gives Since and , we have , hence In both cases, which is the claimed bound.
The posterior bounds in Lemmas 19, 20, 21, 22, and 23 separate the survival of the terminal depth from the normalization accumulated along the observed prefix. The mature-window and initial-window cases cover prefixes on the two sides of the terminal depth , the crude bound supplies a uniform starting point, and the refined statement packages these ingredients into the terminal-posterior control used below.
Let , let , and let . Assume:
(Time scale.) .
(Overlap exponent.) .
(Radius constant.) .
Write , let , and consider the finite signed-depth model of Definition 18 at horizon . For an action/reward-symbol prefix let be its prefix mass under this model, and define the rare-action-history factor from Definition 25 by with the empty product equal to one. Then where the sum ranges over all action/reward-symbol prefixes of length .
⊢ LeanLet Reindex each observed prefix by its action sequence and reward-symbol sequence . This bijection gives
For each action sequence , the product is either or . Hence , including the empty-product case , where . Therefore
After summing over the reward symbols with the actions fixed, the observed prefix mass factors as the behavior action probability product: This is the reward-summed prefix identity for the signed-depth family of Definition 18. Thus
It remains to evaluate the action sum. Since the behavior action law is product-form, For , the inner sum is ; for , it is . Hence
Combining the preceding displays and recalling gives as claimed.
∎Lemma 24 gives the exact second moment of the visible rare-action factor. The displayed expression isolates the contribution of the forced terminal-depth action pattern, and this is the term that enters the observed KL calculation after summing conditional mean differences.
Let and . Write , , , and . Suppose that
(Positive scale.) , , , and .
(Uniform terminal posterior.) For every horizon , every alternative index , every time , and every observed prefix of actions and two-symbol rewards before , the finite conditional terminal-depth probability from Definition 25 satisfies
Then the finite observed-word laws of the signed-depth alternatives of Definition 18, taken before the reward symbols are decoded as real rewards, satisfy
⊢ LeanLet Since , lies in , and hence Also , , and . We prove the claimed bound by induction on the horizon .
Step 1. The zero-horizon case. For , the observed word has a single possible value. The finite-word chain rule gives so the desired inequality is immediate.
Step 2. Full support and projectivity. Fix and assume the bound at horizon . Every observed word has positive mass under each signed-depth alternative. Indeed, the observed coordinate has only one value, and one may choose a hidden trajectory with the displayed observed actions and reward symbols. The initial weights are positive because and ; the behavior weights are positive because ; and the selected hidden transitions have positive reset or continuation weight, with the case handled directly by the reset transition. Thus Moreover, summing the chronological law over the last observed action-reward symbol gives the prefix law: Consequently the last-coordinate chain rule applies and yields where is the conditional law of the next full observed symbol after prefix .
Step 3. The one-step increment. For a prefix of length , write for its action/reward-symbol projection and define For , let be the image of under the bijection The KL divergence is unchanged by this relabelling, so For and reward symbol , write if and if . Let with the conditional ratio evaluated on the finite prefix mass; full support from Step 2 makes this ordinary conditional expectation on the present prefix. Summing the chronological law over all hidden states at time gives, for every , where and . The second equality uses the behavior policy in Definition 18, which chooses the next action with mass independently of the alternative after the observed prefix. The last equality is the two-symbol reward law determined by its mean: for a variable taking values , mass is assigned to and mass to . The positivity assumptions give , so all displayed masses are positive.
By Lemma 30, these means are Using the preceding mass formula and writing out the finite KL sum, Here , and the positive common factor cancels inside the logarithm. The standard two-point comparison, valid for and , yields
Step 4. The conditional chi-square calibration. Set The signed-depth factors are nonnegative: , , and . The action factor also obeys Hence The terminal posteriors are probabilities, so Together with the assumed posterior bound, this gives Substituting the conditional-mean identities from Step 3, The denominator is compared using both and : Since , the lower bound is positive. Applying the preceding numerator and denominator comparisons to the one-step KL bound from Step 3 gives
Step 5. Averaging over prefixes. For an action/reward-symbol prefix of length , define its finite prefix mass using the horizon model by For every observed word of length , the observed state coordinate is fixed, and the mass of under the horizon- observed law is Thus the average increment is bounded by The map is a bijection from length- observed words to length- action/reward-symbol prefixes, because the observed state coordinate has the unique value. Reindexing the finite sum gives By Lemma 24, with , The index set contains exactly the final positions of the length- prefix; if it is empty, and otherwise the shift , , is a bijection. Therefore because and . Hence the increment is bounded by
Step 6. Closing the induction. Using the induction hypothesis and the increment bound, Substituting the definition of gives the asserted bound for , completing the induction.
∎Let and write , , . Then there exists such that
⊢ LeanLet We prove the stated uniform bound by constructing the constant explicitly.
Since , one has . Hence and therefore because is equivalent to , or .
The sequence is bounded above. To see this in the same normalization, set Then and . Writing the first factor tends to by the ratio test, since while . Thus , and in particular there is a finite such that
Define Then . Fix and put From and , we have For every , Indeed, the first inequality follows after multiplication by , since and the second is the standard bound with . Applying this with and raising to the -th power gives
Finally, so monotonicity of the exponential yields Combining the preceding displays, Since was arbitrary, this proves the claim.
Let , let , let , and consider the finite signed-depth model of Definition 18. For an observed prefix , write for the unnormalized filter mass of terminal hidden state , write and for its depth and sign coordinates, and set Assume:
(Positive initial scale.) .
(Positive overlap exponent.) .
(Nontrivial overlap radius.) .
With , define the depth- action-history factor Then
⊢ LeanPut and . For a prefix and a depth , define and with the empty product equal to one. Thus .
We first prove the depthwise invariant For , the initial mass at depth and sign is Summing over and gives since when the prefix is empty.
Suppose the invariant holds for every prefix of length , and let . Write for its first observations and . At reset depth , direct summation of the two reset signs in Definition 18 gives, for each previous hidden state , After multiplying by the previous unnormalized filter mass, the behavior-policy weight, and the reward-emission weight, and then summing over , this yields Because , the invariant holds at depth .
For a positive target depth, write it as with . The last observation of is , and the preceding prefix is the restriction of to its first coordinates. Let be the hidden-transition weight from to under action , for and . The raw filter recursion gives where is the behavior-policy probability of the observed action and . By the transition weights in Definition 18, reaching the positive depth is possible exactly by advancing from depth , and Because , the reward factor on the surviving terms is . Hence
With the same notation, the signed advancing aggregate is obtained by inserting the new sign in the inner sum: Again only survives. For that depth, the transition weights in Definition 18 give Indeed, action preserves the old sign on an advancing transition, while action assigns equal total advancing weight to the two possible new signs. Using once more that the reward factor is at depth , we obtain
The action-history factor has the corresponding one-step recursion. From the definition of , The identities and show that the displayed product is the product defining , followed by the last factor . Thus with the same empty-product convention as before.
Applying the induction hypothesis at depth to the preceding prefix gives The unsigned recursion and the action-factor recursion give which is the same quantity. Equivalently, splitting on the last action, when the common factor is times the induction hypothesis, and when both sides are zero. This completes the induction.
Finally take and . Reindexing the hidden-state sum by its depth and sign coordinates gives The depthwise invariant at , together with , therefore gives as claimed.
Let , let , and consider the finite signed-depth model of Definition 18 with horizon , terminal depth , initial weights , behavior policy , and reward-symbol/hidden-state kernel . Suppose that
(Positive time scale.) .
(Positive behavior perturbation.) .
(Overlap radius.) .
Write where is the signed-depth mixing parameter, and write , . For any observed action/reward-symbol prefix , let denote the unnormalized signed-depth filter mass, after the prefix , assigned to observed state and packed hidden state . Decode the packed hidden state by Then where the outer sum is over all finite signed-depth trajectories of horizon , and is their chronological path weight under , , and .
⊢ Lean1. Let be the set of joint-state prefixes For such a prefix, write for the chronological weight of realizing the observed prefix and ending with hidden state . Summing first over all continuations beyond the observed prefix gives Indeed, partition trajectories by their first observed action–reward symbols and by the joint-state prefix through time . The chronological product factors into , the next behavior probability, and the next kernel probability; summing over the future hidden state and then over gives the displayed identity.
2. For every packed hidden state , the signed-depth reward kernel in Definition 18 satisfies To see this, use the product form of the kernel: the reward-symbol weight depends on and , while the hidden-transition weights sum to one for each current state and action. Hence the left-hand side reduces to The behavior probabilities sum to one, and since , , This proves the displayed one-step payoff identity.
3. Substituting the one-step identity into the prefix decomposition gives The observed state alphabet has one element, so each terminal joint state is uniquely of the form ; this is the only bookkeeping hidden in the notation above.
4. By definition of the unnormalized filter mass, Therefore This is just regrouping the preceding prefix sum by its terminal packed hidden state.
5. Combining the last two displays yields as claimed.
∎Let with , let , let , and consider the finite signed-depth model of Definition 18 at horizon . For an action/reward-symbol prefix of length , define the joint prefix-and-terminal-depth mass at the cut after the first epochs by For each hidden coordinate , let , and let be the unnormalized filter mass, as in Definition 25, assigned after observing to the cut state with hidden coordinate . Then
⊢ Lean1. Let denote the set of joint-state prefixes For such a prefix, write for the real chronological weight of the prefix whose observed action/reward word is . Summing the full trajectory law over all continuations after the cut gives This is exactly the fixed-prefix expansion of the terminal-depth event at the cut after observed epochs; the hypothesis ensures that this cut is a valid time in the horizon .
2. For a joint state , define Then Indeed, interchange the two finite sums. For each fixed prefix , the inner sum over has exactly one nonzero contribution, namely , and therefore returns the displayed indicator and weight.
3. Combining the preceding two displays yields The observed-state coordinate has the singleton value , so this sum is For the signed-depth filter at the cut after the word , the unnormalized mass assigned to hidden coordinate is the terminal-state prefix mass at the joint state : Thus for every , and the claimed identity follows.
∎Let , let , and let be a Boolean alternative, written through Assume
(Positive baseline.) .
(Positive policy scale.) .
(Overlap radius.) .
For every epoch and every observed prefix , the conditional next-reward mean satisfies where the left-hand side is the totalized ratio defining the finite-prefix conditional reward mean, , , and and are the rare-action-history factor and terminal posterior of Definition 25.
⊢ LeanWrite the epoch as a finite ordinal , and set Then . Thus it is enough to prove the identity at the canonical cut after an observed prefix of length , with at least the next epoch present.
Fix such a canonical horizon and write the observed prefix as , where , , and the real reward map is Let be the finite trajectory space with horizon . Its elements record the initial hidden state, the behavior actions, reward symbols, and subsequent hidden states through the remaining epochs of the finite signed-depth weights displayed below. For every , the hidden coordinates are . When , these weights are the signed-depth construction of Definition 18; when , the same formulas are read on the one-point depth set , so only the terminal-depth reset line can occur. For , define and let be the product of the signed-depth initial hidden-state weight, the behavior action probabilities, and the reward-symbol/hidden-transition weights along . Thus, if are the hidden states in , where is the initial hidden-state law with weights , , and is the following one-step reward-symbol and next-hidden-state weight. If , , and , , then where ; for , and with all other hidden transitions assigned weight zero. These formulas are valid for all ; if , only the reset line occurs. Summing the two reward weights gives one, and the displayed hidden-transition weights sum to one for each hidden state and action. Hence The numerator of the totalized finite-prefix conditional reward mean is For a hidden prefix , define its observed-prefix weight by with value when . The unnormalized filter mass at the cut is where and are the depth and sign coordinates of the hidden state .
We next decompose the finite trajectories at the cut. For a hidden state and integer , set with . Since and the displayed normalization of holds at each state-action pair, induction gives Splitting according to the hidden prefix , the next action , the next reward symbol , and the next hidden state , then summing the remaining suffix by , yields This is a partition of the same finite set of trajectories: the data give the first epochs, and the suffix sum runs over all chronological continuations from .
For fixed and , the displayed one-step formula factors into the reward weight and a hidden-transition probability. Hence Using , , and , the inner one-step payoff is Therefore This is the numerator identity supplied by Lemma 28; the displayed derivation records the finite-trajectory regrouping used at the canonical cut. The denominator and the terminal-depth quotient are identified below.
Lemma 27 applies under , , and , and gives the prefix-level identity where, for the length- prefix , At the canonical cut, the time-indexed factor of Definition 25 is the same quantity: since . Thus The empty-prefix case is included by the empty-product convention in Definition 25.
It remains to identify the denominator and terminal-depth numerator. For a trajectory , let be the hidden state after the first observed epochs. In the one-based indexing used for the signed-depth path statement, this same cut state is ; hence the event is exactly the event written there as . Define the finite prefix mass and the terminal-depth prefix mass by and The same cut decomposition as above, now without weighting by the next reward, gives Indeed, the first sum partitions trajectories by their length- hidden prefix, next action, next reward symbol, and next hidden state; the continuation has total mass , and the next one-step mass sums to one. Carrying the terminal-depth indicator at the cut through the same partition gives Thus is the mass of the observed-prefix event and is the mass of the same event together with terminal depth at the cut. The latter identity is exactly the terminal-prefix mass identification in Lemma 29 for the canonical horizon.
The finite-prefix conditional reward mean is the totalized quotient where real division is totalized by . The terminal posterior in Definition 25 is read at this finite prefix through the same totalized convention: If , then because the terminal-depth prefix event is contained in the prefix event; for , these quotients are the usual conditional averages. Combining the numerator identity, Lemma 27, and , we obtain This is the claimed identity for the original and .
Lemma 26 turns the finite-prefix normalization in Lemma 23 into a constant depending on the mixing scale. Lemmas 27, 28, and 29 connect the signed terminal mass, the next reward numerator, and the prefix mass at terminal depth. These identities convert the latent terminal sign into the observable conditional reward mean in Lemma 30, while Lemma 25 converts the resulting terminal-posterior control into a finite observed-word divergence bound. The product reflects the joint requirement that terminal depth survives and that the visible action history carries the rare sequence needed to reveal it.
The likelihood calculation next fixes the observed signed-depth law and combines the exact observed reward mean with a bound on the divergence between the two signs.
For , , and two-point alternative index , define to be the observed-path law obtained by applying the observed-coordinate projection to the trajectory law of the embedded finite signed-depth model with horizon , parameters , and alternative index .
⊢ LeanThe law in Definition 28 is the distribution used in the two-point information comparison. It records the observable image of the embedded signed-depth experiment, so the divergence calculation is stated on the same data observed by an estimator.
Fix , and write . There exists a constant , depending on , such that whenever the following conditions hold:
(Overlap exponent.) , and .
(Reset amplitude.) , and .
(Sequential ignorability.) For every horizon , terminal depth , and alternative , the signed-depth experiment in Definition 18 satisfies Assumption 2.
(Stationary start.) For every horizon , terminal depth , and alternative , the signed-depth experiment in Definition 18 satisfies Assumption 3.
there exists a constant , depending on , such that for every and every , with denoting the observed-path law of the signed-depth alternative , Moreover, for each , each , and each observed prefix , the conditional reward mean satisfies the exact finite-prefix identity where , and and are the rare-action-history factor and terminal-depth posterior from Definition 25. The posterior satisfies The same observed-path laws also satisfy
⊢ LeanFix , put , and set .
1. By Lemma 26, there is a constant , depending only on , such that for every , Fix , , and the two stated model-class hypotheses. Let Since , we have , hence , and therefore .
2. Fix and . The refined posterior estimate in Lemma 23, combined with the choice of in step 1, gives the uniform terminal-depth posterior bound for every auxiliary horizon , every , every , and every observed prefix . In particular, for the present horizon ,
3. Applying Lemma 25 with the posterior bound from step 2 yields the finite observed-word inequality The observed-path law is obtained from the finite observed-word law by the deterministic reward decoding map in Definition 18. The data-processing inequality for this embedding gives Equivalently, by the definition of ,
4. Since , the reset amplitude satisfies Also , so , and . Hence Monotonicity of on this real inequality gives
5. Finally, Lemma 30 supplies, for each , each , and each observed prefix , the exact identity Together with the posterior bound in step 2 and the sharper KL estimate from step 3, this gives all three asserted conclusions: and
∎The identity in Lemma 31 gives the exact observed conditional mean in terms of a rare visible action run and a terminal-depth posterior. The KL bound then combines the survival factor with the policy-overlap factor , matching the fixed-overlap exponent used in the minimax comparison. This is the point at which latent stationarity and observed histories enter the lower-bound calculation together.
The upper-bound argument also uses the stationary mean of each PHIW score. With the observed law already fixed in Definition 26, the next statement bridges the observable score to the target transition operator that controls the truncation term.
Let , let , and let be a raw POMDP experiment. Suppose:
(Model class.) , as in Definition 10.
(Mature time.) and .
Then the depth- score of Definition 16, evaluated under the observed trajectory law of , satisfies Here is the stationary law selected for the behavior-policy transition matrix and is the -fold iterate of the target-policy transition operator applied to the target-policy reward regression
⊢ LeanWrite for the joint state at zero-based epoch , on the same time scale as . For a policy and a function on joint states, set Also write the observable ratio, with its zero-cell convention, as By Definition 10, the model supplies Assumptions 1, 2, 3, 4, and 5, with . Since , .
Let The hypothesis makes a valid epoch, and . We first compute the mean of the score from the state at the beginning of its ratio window: Indeed, the score is a function of the observed trajectory, so its observed-law integral is the corresponding full-trajectory integral, For any bounded and any epoch , Assumptions 2, 1, and 5 give The same conditional calculation at the terminal reward gives The zero-cell convention is compatible with the displayed equalities because and imply . Iterating the preceding identities backward from to , with boundedness supplied by Assumption 4 and finite state spaces from Definition 10, yields the displayed score identity.
It remains to identify the law of . For every function on joint states, At the initial epoch this is Assumption 3. If the identity holds at epoch , then Assumptions 2 and 1 give The selected behavior stationary law is stationary for , so for every , and therefore Induction gives the claim at .
Applying the last identity with and substituting into the score identity gives which is the asserted equality.
∎Lemma 32 states the exact stationary target of each mature score. Under stationarity, a length- observable reweighting transports the behavior stationary distribution through target-policy transitions, so the remaining discrepancy is the contraction-controlled distance between this finite-lag quantity and the stationary target value.
The bounded-reward normalization also pins down the range of the target value. This range fact is used when risks are compared over estimators whose outputs are clipped to the same unit interval.
Let , let , and let be a model in Definition 10. Assume the model satisfies the stationary-start, sequential-ignorability, POMDP-kernel, bounded-reward, latent-stationary-overlap, and policy-overlap conditions of Assumptions 3, 2, 1, 4, 7, and 5. Then the stationary target-policy mean reward of Definition 12 lies in the unit interval:
⊢ LeanLet and be the joint-state transition kernels induced by the behavior and target policies, and let and be their selected stationary laws. For a joint state and action , define the one-step reward mean We also use the local notation which is the target-policy reward regression appearing inside . With this notation, Definition 12 gives
Since , consider epoch . Let denote the unique empty pre-history value at epoch , and define the finite state-history and state-action-history spaces For , write . Let For every and every measurable , define the row selected by the whole state-action history view as By the probability-kernel part of Assumption 1, is a probability measure for every . Define the total function for all ; bounds for will be asserted only on the full-measure set obtained below. Thus, when , one has .
By Assumption 4, holds almost surely under the trajectory law, so the event has full mass under the law of . After both sides are pushed forward to the epoch- conditioning view, Assumption 1 gives the composition-product identity The section argument is now the standard disintegration fact for a composition product: if a set has full mass under the displayed law, then for -almost every its section has full mass under the row . Applying this to gives For every such , the row is a probability measure and the coordinate function is bounded by in absolute value almost surely under that row. The usual norm bound for the integral of an almost-surely bounded real function under a probability measure therefore gives Equivalently, using the definition of ,
Fix and with and . Put The singleton mass of is obtained from three separate ingredients, none of which gives it alone. The first is Assumption 2, which says only that the law of factors as the law of followed by the behavior kernel , so that The second ingredient is the epoch- history projection: the pre-history coordinate at epoch is forced to be , so the event is exactly . The third is Assumption 3, which identifies the law of and gives Combining the last two displays with the two strict-positivity hypotheses yields If a full--measure set omitted this singleton, its complement would have positive -mass. Hence the almost-sure bound from the previous step holds at , and therefore
The selected behavior and target stationary laws are probability vectors: by Definition 6 each lies in , the simplex of nonnegative probability vectors on the joint-state alphabet, fixed by the corresponding policy-induced kernel. Unpacking this simplex membership gives Now suppose and . By Assumption 7, If were not positive, its nonnegativity would force , and substituting that into the displayed overlap inequality would give , a contradiction. This uses the displayed overlap inequality, the nonnegativity , and the local hypothesis ; the radius condition plays no further role in this step. Thus .
The action-coordinate argument uses the local hypothesis , the nonnegativity of the behavior carrier, and the policy-overlap inequality. The inequality comes from Assumption 5, which for the policy-overlap constant gives the nonnegativity comes from Assumption 2, which is where is declared a probability vector. If were not positive, that nonnegativity would force , whence , a contradiction. Thus . The row bound from the previous step applies, and hence
For any with , the target reward regression satisfies The equality in the third line uses ; the last inequality uses the preceding step when , while the corresponding summand is zero when ; and the final equality is the target-policy probability-vector property in Assumption 5. Two distinct facts meet in that last equality, and they come from different places. Which set the sum runs over is a carrier fact: Definition 3 fixes the action coordinate as the Boolean two-point alphabet identified with , and it imposes nothing else — in particular no nonnegativity and no normalization. That the two values at are nonnegative and add to one is the separate probability-vector restriction on , which is imposed in Assumption 5. So the normalization is the sum of exactly two action probabilities at , and no larger action space is involved.
Finally, Here gives the third line; in the fourth line the summand is zero when , and otherwise the regression bound from the previous step applies. Thus which is exactly membership of in the closed real interval .
Lemma 33 records the bounded target range implied by the model-class reward normalization and the target stationary distribution. This range is the common scale for the clipped PHIW estimator and for the bounded observable estimators used in the minimax comparison.
The radius-sensitive upper calculation refines the estimator bound by keeping the latent-overlap distance explicit.
For every trajectory length , observable alphabet size , latent alphabet size , and history depth , let Suppose that
(Mixing scale.) .
(Overlap scale.) .
(Radius parameter.) , with as in Definition 19.
(Trajectory length.) .
(Model class.) The experiment belongs to in the sense of Definition 10.
(History depth.) .
Then the clipped partial-history importance-weighted estimator from Definition 16, evaluated under the observed trajectory law induced by , satisfies where is the stationary target-policy mean reward from Definition 12.
⊢ LeanWrite for the observed trajectory law induced by , and let be the overlap-distance radius from Definition 19. Define the unclipped partial-history average with as in Definition 16. Since and , one has , so . The conditions defining in Definition 10 give and . The finite observation space makes each score measurable. Moreover , and the sequential-ignorability, policy-overlap, and bounded-reward restrictions in Definition 10 imply the window envelope for every . Thus each score is square-integrable under , and so is . Moreover, Lemma 33 applies to the present horizon and model class membership, giving
For every and every , Applying this pointwise with and , and then using the usual bias–variance identity for the square-integrable random variable , yields
We next bound the bias term. Let and be the stationary laws of the behavior- and target-policy transition kernels, and let be the target-policy reward regression appearing in Definition 12. Define On states with , the horizon condition and the model-class restrictions give ; on states with , the definition gives . Hence for every , and therefore The kernel and policy restrictions in Definition 10 make a Markov transition kernel, and Assumption 6 supplies for all probability laws on . Consequently, for any bounded function and any states , the dual total-variation inequality gives Iterating this one-step oscillation contraction yields Applying the last display to gives By Lemma 32, every score in the average defining has mean Therefore For each state with , the functions and agree on the behavior support, so Lemma 11, applied with the target policy and initial state , gives For states with , the common prefactor makes both weighted summands vanish, regardless of the two function values. Hence By Assumption 7, included among the restrictions in Definition 10, . Thus implies , so By the meaning of the target-policy stationary law used in Definition 12, is invariant for . Thus, for every function and every integer , Applying this identity with and yields Combining the preceding displays, For any function with oscillation at most , It remains to prove the stationary-law distance subclaim This subclaim uses the pointwise comparison from Assumption 7. Throughout this step, for probability laws on the finite joint-state space. With this convention, any two probability laws have total-variation distance at most , since If , then for every , and the two probability laws have equal total mass, so . Hence If , define The latent stationary overlap condition gives for every , and so is a probability law. The identity then gives Therefore Thus in all cases,
The variance estimate supplied by Lemma 14, with and , gives
Since and , Substituting this and the variance bound into the clipped bias–variance inequality proves
In Lemma 34, the first term is the squared truncation contribution after observable lags, scaled by the radius . The second term is the sampling contribution from estimating with length- products of observable policy ratios. Choosing the radius-calibrated depth in Definition 20 balances these terms in the shrinking-overlap results.
The matching radius-explicit KL statement records the lower-bound calculation with the same radius normalization.
Let . Define Then there exists a constant such that, for every , , , and with , if and if is the radius in Definition 19, then and the observed-path laws and of the signed-depth alternatives satisfy
⊢ LeanFix .
1. Apply Lemma 31 with this value of . It gives a constant , and it remains only to supply the sequential-ignorability and stationary-start hypotheses for the signed-depth experiments over the range requested there. For every auxiliary horizon , terminal depth , and sign alternative , the finite construction underlying Definition 18 has the observed-state behavior kernel and the embedding into the observable-data experiment preserves this factorization, so Assumption 2 holds. The same construction starts the joint state from its behavior stationary law, so Assumption 3 holds for the same , including the zero-horizon case. Hence Lemma 31 supplies, for every , , , and , the sharpened observed-path bound This KL estimate is one conjunct of the conclusion; the radius comparison is verified separately.
2. It remains to prove Since , we have and . First suppose . Then For the lower bound, is which follows from , , and . For the upper bound, is Multiplying by , this becomes and this follows from and .
3. Now suppose . Then For the lower bound, is equivalently , which is immediate. For the upper bound, is Multiplying by , this becomes , equivalently , which follows from .
Combining the two cases gives the asserted radius comparison, and the KL bound is the one obtained in Step 1 with and .
∎The inequalities in Lemma 35 connect the construction amplitude to the normalized radius from Definition 19. Consequently, the lower calculation can be stated on the same radius scale as the estimator bound in Lemma 34.
The final group records the parametric two-experiment comparison used to anchor the risk at the usual scale across all radii in the model class. The first two statements compute the separation in target value and the observed KL budget for the singleton-state pair; the following bounded-risk and information inequalities convert those calculations into risk floors.
Let , let , and define the sign Put with value when . Suppose the scalar parameters satisfy:
(Initial scale.) .
(Overlap scale.) .
(Radius scale.) .
Let be the singleton-state finite experiment with one observed state and one hidden state, fair behavior and target policies , reward support , and reward-transition kernel that assigns, from every state-action pair, probability to reward and next state . Then the stationary target-policy mean reward of Definition 12, evaluated in this experiment, satisfies
⊢ LeanLet , viewed first as the finite reward model before applying the embedding into the raw experiment. Write for the target-policy stationary law on the singleton joint-state space associated with this finite model. The construction of , under , , and , belongs to ; in particular the target stationary law is a probability vector. Thus The equality between the finite encoding and the raw experiment preserves this stationary law, so the same is the law used in as defined in Definition 12.
For every joint state and action , the conditional reward mean in the finite kernel is The displayed calculation is exactly the reward-kernel mean, with the encoding and .
By Definition 12 and the finite-to-raw value identity for this encoding, Substituting the preceding reward-mean identity gives The target policy in is fair, hence for every , Therefore using the total mass identity for .
∎Let satisfy , and set . Let and be the two finite singleton-state experiments with one observed state , one hidden state, initial mass at the single joint state, fair behavior and target policies on , and reward code decoded as . Under , for , the reward-transition kernel returns to the single joint state and assigns reward probability Writing for the image of the decoded trajectory law under the observed-path projection, the required KL budget holds in the direction
⊢ LeanWrite and for the two finite reward-code models with Boolean indices and , so that the associated signs are and . Let Since , we have and .
The two finite models use the same reward decoder . Therefore applying the decoder and then projecting to the observed path is a common measurable map, and KL contracts under this map: where denotes the finite observed-path law before decoding the reward code into .
For one epoch define probability laws and on by The first factor is the fair action law, and the second factor is the reward-code law. Because , every displayed cell mass is positive.
Multiplying the singleton initial mass, the fair action probabilities, and the displayed time-homogeneous reward-transition probabilities gives the product form The singleton observed state contributes unit mass at each time, so the only nontrivial coordinates in each factor are the fair action and the reward code.
The one-epoch divergence is bounded by Let be the fair law on , let be the law on assigning mass to reward code , and set Since , both and lie in . The law is obtained from by the bijective coordinate map . Hence invariance under this reordering and the common action coordinate give The Boolean Bernoulli laws are images of the two-point real laws under the measurable map that records whether the point is , so data processing gives the upper bound by the corresponding real two-point divergence. Evaluating that finite divergence yields For , the function has value at and derivative Thus on this interval, and substituting gives the displayed one-epoch bound.
Since every one-epoch cell under has positive mass, ; on this finite space the log-likelihood ratio is integrable. Tensorization of KL for the -fold product laws gives two finite-valued conclusions: the product divergence is finite, and after applying the usual real value map to the extended nonnegative divergence, The one-epoch bound from the previous step is finite, so the same real-value comparison gives The product representation of the finite observed-path law therefore yields Because the displayed product divergence is finite, comparison with any finite numerical bound is equivalent to comparison after taking this real value.
Finally, using and . The finite-valued comparison from the preceding step gives
The contraction inequality from the first step and the finite observed-path bound from the last step give as claimed.
∎Before applying the two-point inequality, the next statement records a uniform boundedness fact for clipped observable-data estimators. It uses the target-value range from Lemma 33 and the estimator range built into Definition 14.
Let and let . Let be an observable-data estimator of the admissible class used in Definition 14: for each observed-alphabet size and policy pair, is a measurable real-valued function of the observed path taking values in . For every as in Definition 10, with evaluated at the observed-alphabet size and policies of , the squared risk for the target value of Definition 12 satisfies
⊢ LeanFix and write an observed path as For this path set The admissibility condition in Definition 14 gives and Lemma 33 gives Hence Equivalently, Multiplying these two nonnegative quantities yields so
The function is measurable because is measurable by admissibility and the target value is constant in the observed path. The observed-path law is a probability law, and the pointwise bound from the previous step gives integrability together with By the squared-risk definition in Definition 14, the left-hand side is exactly Therefore
Lemma 38 places all admissible estimator losses on a common bounded scale. This supplies the integrability and range control used when the two-point comparison is embedded into the minimax risk.
Let , , and let as in Definition 10. After identifying their common observed alphabet, write for the observed-data law of and for the target value of from Definition 12, . Let be a nonnegative separation radius and let be a nonnegative KL upper bound. Assume:
(Common observable structure.) The two models have the same observed-alphabet cardinality, the same behavior policy, and the same target policy.
(Information bound.) ,
(Target separation.)
Then every raw observable-data estimator , viewed as a real-valued measurable function of the observed-alphabet size, the policy pair, and the observed path, with integrable under at for , satisfies Consequently, the minimax squared-error risk of Definition 14 satisfies
⊢ LeanWrite the common observed alphabet size as , and use the common observable structure to identify the behavior and target policy carriers of the two models. Put where is the target value from Definition 12. The information bound and give where denotes the real value of the finite extended-real KL divergence. Hence By the Bretagnolle–Huber affinity inequality, using and finite KL divergence, Combining the two displays yields
Fix a raw observable-data estimator satisfying the measurability and integrability hypotheses. With the common observable inputs fixed, define and define the two error events The separation condition and the two-point testing bound for real-valued parameters give
Since , for each , The squared-error integrability assumptions therefore allow the elementary tail-to-mean comparison
Combining Equations 1 and 2 gives If , then the last maximum is , and Equation 3 gives If , the same argument gives These two cases prove
It remains to pass from the fixed two models to the minimax risk in Definition 14. The observable-estimator class is nonempty, for example it contains the constant-zero estimator. We first record the bounded-risk input needed for the minimax reduction. Fix an admissible observable estimator . For every model , with observed alphabet size , behavior carrier , target carrier , observed law , and target value , Lemma 38 gives Consequently, which is the boundedness hypothesis needed in the two-point minimax reduction. For such an estimator, the two integrability requirements used above hold at and . These are two separate statements, one for each model, and each is taken at that model’s own observable inputs before any identification of carriers: at the estimator is measurable by admissibility and its range condition gives so the fixed-model squared error is dominated by the constant , since Only after both integrability statements are in hand do the common-policy hypotheses and rewrite the two carriers into the single pair used for above, so that the two integrals are integrals of one function against the two observed laws. Applying Equation 4 to each admissible observable estimator and then the standard two-point reduction to the minimax value yields which is the claimed minimax lower bound.
∎Lemma 36 sets the value separation for the singleton-state Rademacher pair, and Lemma 37 gives the corresponding observed KL budget. The bounded-risk statement in Lemma 38 fixes the loss scale for admissible clipped estimators, and the generic reduction in Lemma 39 then translates a common-policy, low-divergence pair with target separation into a lower bound for every observable-data estimator.
Fix a positive mixing scale and a positive log-overlap scale . For every horizon with and every radius , let and be the two raw POMDP experiments generated by the singleton-state, singleton-observation Rademacher construction with rewards in , fair behavior and target policies , and the two opposite signs of the construction’s parametric reward tilt. Then:
(Class membership.) Both and belong to the latent-overlap model class of Definition 10.
(Common policies.) Each experiment has equal behavior and target policies, and the two experiments share these policies:
(Uniform minimax floor.) The minimax squared-error risk of Definition 14 satisfies
(Pointwise two-experiment floor.) For every observable-data estimator , viewed as a measurable function of the revealed alphabet and the policies, under the usual squared-error integrability conditions for the observed trajectory laws of and , where is the stationary target-policy mean reward of Definition 12.
Fix and put Since , . For each sign , define a singleton-state experiment with , , fair policies and reward-transition kernel Let be the experiment with , and the experiment with . The weights above are nonnegative and sum to one. The hypotheses give the positive parameter-scale requirements and appearing in Definition 10. The Markov-kernel, sequential-ignorability, stationary-start, and bounded-reward requirements in Assumptions 1, 2, 3, and 4 hold by this generated singleton construction. The policy-overlap condition in Assumption 5 holds because and . For either policy , for all probability laws on the joint state space. Put Since , , and therefore . Thus which is the uniform-contraction requirement in Assumption 6. Finally, the latent stationary overlap condition in Assumption 7 is immediate: , and hence for every state whenever . Thus both alternatives belong to as defined in Definition 10. The same displayed definition of the policies also gives
The singleton alternatives just constructed are the two alternatives covered by Lemma 36. With the sign convention of Step 1, that result gives Consequently For the divergence calculation, let be the reward code and set Define the one-epoch coded observed law by Since , every displayed cell has positive mass. Thus , and the log likelihood ratio is integrable. If and denote the reward-code marginals of and , then The fair action factor is common to the two signs, so the product rule for KL with a common second factor gives Writing out the two Bernoulli masses gives For , using and . Applying this with yields the one-epoch bound Let The generated finite trajectory construction has this product as its coded observed-word law. By finite-product KL tensorization, with the absolute continuity and integrability supplied by the positive one-epoch masses, Because , , and , Finally, let be the decoding map that replaces each reward code by and keeps the observed state and action coordinates. The observed laws in the statement are the pushforwards KL divergence is monotone under measurable pushforward, so the preceding coded product bound gives This is the observed KL budget in Lemma 37. The positive one-epoch masses give , hence also , and the displayed upper bound gives finite KL divergence.
Let be any observable-data estimator satisfying the measurability and two integrability assumptions in the statement, and write where the equality uses the common policies from Step 1. The two-point risk inequality applied here is Lemma 39, used in its raw-estimator form for the estimator . Its structural hypotheses are available here: the observed alphabets both have one state and the behavior and target policies are common by Step 1. Step 2 supplies together with finite KL divergence. The present step supplies the required measurability of for every input , and the two squared-error integrability assumptions for the two displayed risks.
Apply Lemma 39 directly to the raw estimator , the two observed laws the target values and the numerical choices The separation condition is supplied by Step 2, and the absolute-continuity and finite-KL hypotheses are supplied by the final display of Step 2. Hence Finally, which is exactly the pointwise two-experiment floor in the statement.
It remains to record the minimax lower bound. For the supplied radius , the two singleton alternatives with the same kernels and policies satisfy the hypotheses of the minimax conclusion in Lemma 39: Step 1 gives their common observable cardinality, common behavior policy, common target policy, and membership in ; Step 2 gives with finite KL divergence. Applying Lemma 39 with yields where is the observable minimax risk of Definition 14. Since the desired minimax floor follows. Together with the membership and policy identities from Step 1 and the pointwise inequality from Step 3, this proves all four asserted conclusions.
Lemma 40 complements the signed-depth lower construction by giving a radius-uniform benchmark at the parametric scale. The experiments share behavior and target policies, so the displayed floor reflects value uncertainty in the observed rewards themselves and applies throughout the latent-overlap model class.
Verification note
This note records the machine-verification scope of the paper.
What is checked. Every displayed assumption, definition, lemma, proposition, and theorem of the main text and of this appendix is either mapped to a declaration of an accompanying Lean 4 development, listed under Statement map below, or flagged under Presentation-level definitions as notation introduced for the prose alone. Each displayed proof renders the corresponding Lean proof. The Lean development is the verified artifact; the displayed statements and proofs are human-readable transcriptions of it, and the constants, exponents, and inequalities they display are the ones that appear in the machine-checked declarations.
Toolchain and artifact version. The development compiles under the Lean 4 toolchain leanprover/lean4:v4.33.0 against mathlib4 at revision db584cd6d46c92f209a44c0f1c829460d327499d, in the module tree CausalSmith/Stat/STAT_PomdpLatentOverlapMinimax_Research, from the artifact source tree at commit f2e7e979167a2f3f9242e93c680c7ba1667b6dcc, which is the commit recorded in the verification contract accompanying this paper.
What the receipts certify. The receipts behind this note are declaration-level: for each declaration listed below, the recorded check is that it compiles under the stated toolchain and that its source carries no sorry and no admit. That is the exact scope of the trust claim made here. Readers who additionally want the kernel-level axiom footprint of a particular declaration can obtain it directly by running #print axioms on that declaration in the accompanying development.
External inputs. No result in this paper is stated conditionally on a published theorem: the cited literature supplies framing, motivation, and comparison only, and no bibliographic item enters any displayed statement as a hypothesis.
Declarations from outside the module tree. Lemma 5 is carried by Causalean.Mathlib.Probability.CertifiedFiniteMarkovExpectation.stationaryRewardInterval_sound_of_chunked, reached through Helpers/InsulinGrid.lean. It belongs to the same verified library and compiles under the same toolchain.
Statement map. Each anchored object and the declaration it is mapped to, with the file of the module tree that contains it:
Definition 1:
policyFactorinBasic.lean.Definition 2:
FullTrajectoryinBasic.lean.Assumption 1:
PomdpKernelLawinBasic.lean.Definition 3:
PolicyinBasic.lean.Assumption 2:
SequentialIgnorabilityinBasic.lean.Definition 4:
policyKernelinBasic.lean.Definition 5:
JointStateinBasic.lean.Assumption 3:
StationaryStartinBasic.lean.Assumption 4:
BoundedRewardinBasic.lean.Assumption 5:
PolicyOverlapinBasic.lean.Definition 7:
mixingAlphainBasic.lean.Assumption 6:
UniformContractioninBasic.lean.Assumption 7:
LatentStationaryOverlapinBasic.lean.Definition 8:
ObsViewinBasic.lean.Definition 10:
LatentOverlapClassinBasic.lean.Definition 11:
rewardRegressioninBasic.lean.Definition 12:
targetValueinBasic.lean.Definition 14:
minimaxRiskinBasic.lean.Definition 15:
ratioinBasic.lean.Definition 16:
phiwEstimatorinBasic.lean.Definition 17:
historyDepthinBasic.lean.Definition 18:
signedDepthFamilyinHelpers/SignedDepth.lean.Definition 19:
overlapRadiusinBasic.lean.Definition 20:
radiusAdaptiveDepthinBasic.lean.Theorem 1:
uniform_overlap_frontierinTUniformOverlapFrontier.lean.Theorem 2:
fixed_c_minimaxinTFixedCMinimax.lean.Theorem 3:
unit_overlap_boundaryinTUnitOverlapBoundary.lean.Theorem 4:
shrinking_overlap_frontierinTShrinkingOverlapFrontier.lean.Definition 24:
insulinGridinHelpers/InsulinGridCore/Model.lean.Lemma 1:
insulinPolicy_contractioninHelpers/InsulinGridCore/Model.lean.Lemma 2:
insulinStationaryLaw_isStationaryinHelpers/InsulinGridCore/Model.lean.Lemma 3:
insulinTargetRewardRegressioninHelpers/InsulinGridCore/Semantic.lean.Lemma 4:
insulinImmediateRewardRegressioninHelpers/InsulinGridCore/Semantic.lean.Lemma 6:
insulinBiasInterval_boundsinHelpers/InsulinChunkAssembly.lean.Lemma 7:
insulinGrid_bias_certificateinHelpers/InsulinGrid.lean.Theorem 5:
finite_insulin_demonstrationinTFiniteInsulinDemonstration.lean.Definition 25:
filterQuantitiesinHelpers/ObservedFilter.lean.Lemma 8:
integral_sq_phiwScore_leinHelpers/PhiwWindowMoments.lean.Lemma 9:
abs_covariance_phiwScore_le_overlapDecayinHelpers/PhiwWindowMoments.lean.Lemma 10:
integral_phiwScore_mul_eq_futureFunctioninHelpers/PhiwFutureEndpoint.lean.Lemma 11:
markovOperatorIter_congr_on_behaviorSupportinHelpers/StationaryRewardSupport.lean.Lemma 12:
abs_covariance_phiwScore_le_disjoint_sharpinHelpers/PhiwFutureEndpoint.lean.Lemma 13:
sum_abs_covariance_future_leinHelpers/PhiwVariance.lean.Lemma 14:
variance_phiwRaw_leinHelpers/PhiwVariance.lean.Lemma 15:
phiw_upperinHelpers/PhiwUpper.lean.Lemma 16:
signedDepth_uniformContraction_actionOneinHelpers/SignedDepthStationarity.lean.Lemma 17:
signedDepth_stationaryLaw_applyinHelpers/SignedDepthStationarity.lean.Lemma 18:
signedDepth_membershipinHelpers/SignedDepthMembership.lean.Lemma 19:
pow_one_sub_le_signedDepthPosteriorLowerWindowinHelpers/TerminalPosteriorBound.lean.Lemma 20:
signedDepthRawPosterior_le_product_of_leinHelpers/TerminalPosteriorBound.lean.Lemma 21:
signedDepthRawPosterior_le_product_of_ltinHelpers/TerminalPosteriorBound.lean.Lemma 22:
signedDepthRawPosterior_le_ratio_powinHelpers/TerminalPosteriorBound.lean.Lemma 23:
signedDepthRawPosterior_le_refinedinHelpers/TerminalPosteriorBound.lean.Lemma 24:
rareAction_secondMoment_eqinHelpers/SignedDepthObservedChain.lean.Lemma 25:
signedDepth_finiteObs_klDiv_leinHelpers/ObservedKL.lean.Lemma 26:
exists_uniform_signedDepth_bootstrap_factorinHelpers/TerminalPosteriorBound.lean.Lemma 27:
signedDepth_terminalMass_invariantinHelpers/SignedDepthFilterInvariant.lean.Lemma 28:
signedDepth_prefix_nextRewardNumeratorinHelpers/SignedDepthFilterRecursion.lean.Lemma 29:
terminalPrefixMass_eq_sum_signedDepthFilterMassinHelpers/ObservedFilter.lean.Lemma 30:
conditionalRewardMean_signedDepth_allinHelpers/ObservedFilter.lean.Definition 28:
signedDepthObsLawinHelpers/ObservedKL.lean.Lemma 31:
observed_path_klinHelpers/ObservedKL.lean.Lemma 32:
integral_phiwScore_eq_init_targetIterinHelpers/BiasVariance.lean.Lemma 33:
targetValue_mem_unitinHelpers/Kernels.lean.Lemma 34:
radius_sensitive_phiwinHelpers/BiasVariance.lean.Lemma 35:
radius_explicit_observed_klinHelpers/ObservedKL.lean.Lemma 36:
parametricPair_targetValueinHelpers/TwoPoint.lean.Lemma 37:
parametricPair_observed_klDiv_leinHelpers/TwoPoint.lean.Lemma 38:
observedRisk_le_fourinHelpers/Kernels.lean.Lemma 39:
two_point_floor_explicitinHelpers/TwoPoint.lean.Lemma 40:
uniform_parametric_floorinHelpers/TwoPoint.lean.
Presentation-level definitions. The following definitions correspond to no single Lean declaration. They name objects the prose needs in order to display the argument, and which the Lean development carries inline rather than through a named definition; no claim is attached to them beyond the notation they fix: Definition 6, Definition 9, Definition 13, Definition 21, Definition 22, Definition 23, Definition 26, Definition 27.
Proofs of the main results
Write Then , , and .
Step 1. Adaptive-depth calibration. Set There are and , depending only on , such that, for all and all , Indeed, since , enlarge so that If , then by Definition 20, and , hence If , then , , and the same logarithmic bound gives Thus . The floor bounds give Consequently, Taking proves the calibration.
Step 2. Upper bound. Let For , put . By Definition 20, , and Lemma 34 gives, uniformly over , By the definition of , The calibration in Step 1 and the inequality therefore imply Since is the infimum over observable-data estimators in Definition 14, this also yields
Step 3. Constants for the lower comparison. From Lemma 35, choose . With , define Then . Put All these constants are positive and depend only on .
Step 4. Lower bound in the small-radius branch. By Lemma 40, If , then with both sides equal to when . Hence and therefore
Step 5. Lower bound in the large-radius branch. Assume . Then . Define Because , we have , and Let and be the two signed-depth alternatives of depth from Definition 18, and write and for their target values. By Lemma 18, both alternatives belong to . With denoting the terminal depth mass in that construction, the same result gives Define the separation radius The target-value identities give the separation condition needed for the two-point bound, Also, by Lemma 35, and, orienting the two observed laws as against , Since , the choice of gives The two alternatives have the same observed alphabet and the same revealed policies by their construction; the displayed KL bound, in the same orientation, supplies the information and finiteness hypotheses with . Applying Lemma 39 therefore gives The depth mass in Lemma 18 satisfies and the ceiling upper bound implies Thus Combining this hidden lower bound with the parametric floor and using and , we get
Step 6. Assemble the frontier. Set Here uses the same calibration constant fixed in Step 1. Thus and . Since , the two upper bounds from Step 2 with constant also hold with constant . Together, Steps 2, 4, and 5 prove, for all and , together with
Step 7. Immediate estimator in the regime. Let and suppose there is such that, eventually, Define For all sufficiently large , apply Lemma 34 with . Since , This is the asserted eventual parametric bound for the immediate estimator.
∎Set By and , the constants satisfy and, using the definitions in Definitions 7 and 1, The exponent in the theorem can be written as Apply Lemma 15 with . It gives constants and such that, for all , For the signed-depth family, the behavior policy in Definition 18 is a memoryless function of the single observed state, independent of the latent coordinate and the past. Consequently, for every horizon , terminal depth , and sign , the action kernel satisfies Assumption 2. The same alternatives start from the behavior stationary law for every horizon, including the zero-horizon case. Indeed, for the signed-depth constant-policy stationary law at action-one probability has masses and the behavior policy in Definition 18 chooses action one with probability . Hence its stationary law is precisely the initial hidden-state law The chronological path law is generated from this initial distribution, so the time-zero joint-state marginal is this stationary law; when , the same identity is the whole marginal check, and for it agrees with the behavior-stationary masses displayed in Lemma 18. Thus Assumption 3 holds for every horizon. These are the two structural hypotheses needed for Lemma 31. Therefore that lemma gives a constant , depending on , such that for all and , Let and define Since , every factor in this product is positive, so . Finally set Since and , the maximum is positive; moreover the second argument is bounded by the maximum. Thus
Define Fix an integer . Then , , and Indeed, and multiplication by gives the displayed inequality. Put Since and , one has . Moreover, and therefore
The preceding depth choice makes the two signed-depth alternatives statistically close. First, because , , and . Hence Combining this inequality with the KL bound from Lemma 31 gives
Let and be the two signed-depth experiments of Definition 18. They have one observed state and latent states. Their revealed behavior and target policies are common, because the policy part of the construction is independent of the sign . By Lemma 18, both experiments belong to , and their target values are where Set Then , and the two target values satisfy the separation bound The two embedded observed laws have the same finite real-reward support, They assign total mass one to , because the construction has one observed state, Boolean actions, and rewards in . We now check that every atom in this common support has positive mass under both signs. Fix , and choose a fixed latent sign . Consider the full hidden lift that keeps the latent state at depth zero and sign : Since and , Also , so the initial mass of the chosen hidden state is positive: For each action in the observed word, because . Along the displayed hidden lift, the current depth is . Hence the reward factor is and the hidden transition that resets to the same depth-zero sign has weight Thus this single compatible hidden lift has strictly positive finite-path probability for each sign , and summing over all lifts gives The reward decoding used in the embedding is the same for the two signs, so the embedded laws preserve this common atomic support. Therefore any measurable set with zero mass under one embedded observed law contains no atom of , and consequently has zero mass under the other embedded observed law. Hence and the finite KL displayed above is an ordinary KL bound between mutually absolutely continuous observed laws.
Apply Lemma 39 to the present pair with , taking the alternative as the lemma’s first model and the alternative as its second: and . With this assignment the lemma’s divergence hypothesis reads which is exactly the direction bounded above, and its absolute-continuity hypothesis reads , which the preceding step supplies. The common observable structure was checked in the preceding step, the KL bound gives the information hypothesis with , and the displayed target-value calculation gives Therefore every measurable observable estimator , under the stated squared-error integrability conditions for both alternatives, satisfies Commuting the two entries of the maximum, this is exactly The two signed-depth alternatives belong to and share the revealed inputs . For the witness clause of the theorem the two squared-error integrability conditions are hypotheses, imposed on the estimator and discharged by whoever supplies it; nothing above derives them. They are derived only when is further restricted to the admissible carrier of Definition 14, whose members are measurable and uniformly -valued by definition: for such an estimator the squared error at either alternative is a measurable function bounded by on a probability space, hence integrable, so the preceding pointwise lower bound applies to every admissible estimator. Taking the supremum over the model class, then the infimum over these admissible estimators, yields
It remains to compare with . Since and , also . Thus so the denominator in the depth mass is strictly positive and satisfies The numerator is nonnegative, and division by a positive number at most one gives The upper ceiling inequality and the identity give Therefore Squaring and using yields Multiplying by gives Combining this with the two-point floor proves and, for every measurable observable-data estimator satisfying the two integrability conditions,
The upper bound follows from the estimator bound in the first step. Since , By Definition 14, the minimax risk is bounded above by the worst-case risk of any admissible observable estimator, in particular by that of . Thus Together with the lower bound from the preceding step, this proves all asserted inequalities for every , and the chosen supplies the stated signed-depth witnesses with one observed state and latent states.
Write Since and , one has and .
1. Define The preceding inequalities imply . For any , Lemma 34 applied at depth and overlap constant gives, uniformly over , Taking the supremum over the model class yields
2. Let . Write and for the behavior and target stationary laws on the finite joint state space , and define The stationary-law total-variation bound at overlap radius gives and at unit overlap Definition 19 gives Thus Since the displayed total-variation quantity is a half-sum of absolute values, it is nonnegative, so Unfolding the definition of total variation yields For each joint state , the nonnegative summand is bounded by the whole finite sum, and hence Therefore for every , and consequently .
3. Set Then , , and therefore ; also . For every , Lemma 40 at gives which is For the upper bound, the minimax risk is bounded above by the worst-case risk of any admissible estimator with nonnegative squared risk; applying this to and using Step 1 gives
4. Finally, Therefore which is the asserted bound on the hidden-state exponent.
∎Put Since and , one has , , , and hence .
1. A radius-depth calibration. Let Then , There are and , depending only on , such that for every and every , Indeed, choose so that for all . Fix such a , put , and set . The relation and Definition 19 give , hence . If , then by Definition 20, and , so If , then . With the choice of gives , so the cap in Definition 20 is inactive and . The floor bounds imply Since the case is bounded by Taking proves Equation 5.
2. Constants. Apply Theorem 1 to obtain constants , with , and a threshold , all depending only on , such that its two-sided risk comparison and radius-calibrated estimator bound hold for all and . Define and Finally set Then , because , , and .
3. Comparing the two rate surfaces. Fix a deterministic sequence with and . Choose such that Let . For , put . Then , and Since , Raising these inequalities to the positive power gives Thus the theorem’s local rate, evaluated along the fixed sequence , and the fixed-radius rate at satisfy Equivalently, since in the statement denotes this fixed-sequence quantity, Combining Equation 6 with the two-sided comparison from Theorem 1 gives
4. Attainment by the displayed shrinking depth. For , define Then , and Therefore Definition 20 gives where is the depth displayed in the theorem statement. In particular . Applying Equation 5 with yields Now fix . By Lemma 34, the estimator from Definition 16 satisfies Using and the definition of , Together with Equation 9 and , this gives Taking the supremum over proves the asserted estimator bound.
5. The bounded local regime. Assume that there is a constant such that for all sufficiently large . Since , the same tail of ’s satisfies Applying the final assertion of Theorem 1 to the sequence gives a constant such that, for all sufficiently large , This is the claimed immediate-depth conclusion.
∎Fix , a policy , and probability vectors on . We prove the displayed total-variation bound; the formulation follows from .
1. Let be the uniform law on . For , , and , define the non-refresh transition weight where , , , , , and for . The refresh construction in Definition 24 gives, after marginalizing over the reward, Therefore the policy-induced matrix can be written as Equivalently,
2. For each fixed , is a probability vector. Indeed, all factors in are nonnegative. Also, because the rounded Gaussian cells partition the clipped glucose outcome, the diet probabilities sum to , the activity probabilities sum to , and the three indicator factors merely select the shifted memory coordinates. Thus
3. It remains to apply the common-refresh contraction calculation. For any finite state space, any probability vectors , and any stochastic matrix , a kernel of the form satisfies since the refresh term contributes . Hence With , , , and , this gives Multiplying by yields the stated version.
∎Write for the transition matrix induced by the policy on the finite state space . We verify the hypotheses of the finite-state contraction fixed-point criterion and then identify the selected fixed point.
Since is a probability law on for each observed grid state , and since the kernel of is a Markov kernel, the policy mixture is stochastic. Thus, for every probability vector on , is again a probability vector on . The initial law specified for is also a probability vector.
By the stochasticity verified in the preceding step, the map sends the probability simplex into itself. The displayed hypothesis of Lemma 2 gives for all probability vectors . The constant is nonnegative and strictly smaller than . Since the probability simplex over the finite set is complete in total variation distance, the contraction mapping theorem gives a probability vector such that Equivalently,
The selected law is defined as the stationary law chosen for whenever such a stationary probability vector exists. The preceding step supplies that existence, so the selected law is itself a probability vector and satisfies Reading this coordinatewise yields which is the asserted stationarity equation.
Fix and abbreviate By the definition of the target-policy reward regression, the left-hand side of the claimed identity is .
For each , the embedded insulin kernel assigns the reward coordinate according to the current glucose component of . Hence its conditional reward mean is Indeed, Definition 24 makes the reward coordinate a deterministic function of the present state, with the displayed values defining ; the subsequent randomization affects the next state coordinates but not this reward coordinate.
Substituting this identity into the regression gives The target policy in Definition 24 is a probability vector on : if , then and , while if , then and . Therefore and consequently .
This is the asserted identity.
∎Fix and write .
Let be the reward category selected by the current glucose coordinate, and let Thus . In the finite insulin kernel of Definition 24, for each action the next-state mass may be written as , with and the reward coordinate is deterministic at the current-state reward category: Therefore finite-space integration gives
The behavior policy in Definition 24 satisfies so for both . By the zero-cell convention in the definition of , the positive-denominator branch applies and hence
Substituting the two preceding identities into the weighted regression gives The target policy in Definition 24 injects exactly when , so Consequently as claimed.
Fix and . For this pair , write the chunked-recurrence data as These are the block partition, block intervals, and partial-sum intervals attached to this fixed . The chunked recurrence first gives the ordinary interval recurrence Indeed, take arbitrary choices from the interval summands on the left and group them by the disjoint covering blocks . The contribution of block lies in . Starting from the tail value , the inclusions show by backward induction that the sum of the block contributions from through lies in . At this is the full sum, and . If the block list is empty, the same argument is just the terminal condition . Thus the initial row together with these coordinate recurrences is the ordinary finite-iterate certificate with rows used in the remaining argument.
Define the true iterates , for all , by Here is the given initial probability vector, while denotes its -step image under . The initial condition gives . If for every , then and exact interval multiplication give and exact interval summation together with the recurrence from the previous step gives Thus, for , In particular, Since is a probability vector and is stochastic, every is a probability vector.
The terminal interval expectation contains the true reward expectation at time : This follows coordinatewise from and , followed by exact interval products and exact finite interval addition.
The one-step residual of is bounded by . Since , exact interval subtraction gives and hence Summing over yields Because is stationary and both and are probability vectors, Since , this gives
For every state , Therefore
Combining the preceding two enclosures, let From and we obtain By the definition of , this is exactly
All arithmetic below is exact rational interval arithmetic with the operations described in Lemma 5.
The residual sums over the states of Definition 24 evaluate to and
Since , the interval expectations computed from the terminal rows are and
Because , the expansion radius in each case is . Hence and
Taking the interval difference, so The endpoint comparisons are then the rational inequalities and This proves both asserted endpoint bounds.
Fix . Write and for the policy-induced transition matrices on obtained from Lemma 1 by taking and , respectively, and write and for their selected stationary laws.
The behavior and target kernels both satisfy the one-half contraction estimate for all probability vectors on . This is exactly the -form supplied by Lemma 1, instantiated first at and then at . The corresponding selected laws are probability vectors and satisfy by Lemma 2.
Let For , with denoting or , let be the exact rational interval matrix recorded by the finite-grid certificate. Its entrywise enclosure is The same certificate supplies an exact rational initial probability row , interval rows and , and an exact chunked recurrence from the first interval row to the second. In the notation of Lemma 5, its rational inclusions verify For every output state, the recorded blocks partition , every block has at most states, and the block and partial-sum intervals satisfy exactly the inclusion chain in Lemma 5. Because the certificate has zero approximation steps, this verified recurrence links its initial interval row directly to its successor row. Passing from the chunked certificate to the finite recurrence used by that lemma preserves both rows and all the rational inclusions. For the reward certificate, put Since for every grid state, the reward bound in Lemma 5 is . Define After that conversion, the terminal row is and the successor row is , so the residual and stationary-reward interval are The initial-row, recurrence, reward-bound, contraction, and stationarity hypotheses of Lemma 5 are therefore satisfied for both policies after passing each chunked recurrence certificate to its associated finite recurrence certificate. Under this conversion, the terminal interval row is , the successor interval row is , and the stationary-reward interval produced by Lemma 5 is exactly which is the chunked interval used here. Hence
Define the subtraction interval Interval subtraction is sound, so the two inclusions from the previous step give
The reward-regression identities in Lemmas 4 and 3 give, for every , and The finite-grid embedding preserves the policy-induced kernels entry by entry: Consequently the stationary laws in the interval certificate are exactly the behavior- and target-policy stationary laws entering the two regression identities. Combining these kernel identifications, the two regression identities, and Definition 12 yields Therefore
By the exact rational interval evaluation in Lemma 6, the endpoints of the subtraction interval satisfy Since , we have The last two displays imply as claimed.
Fix . We verify the five asserted clauses.
By Definition 24, the observed coordinate has and the latent diet-memory coordinate has Thus the joint-state alphabet has
For a memoryless policy , let be the reward-marginalized joint-state transition matrix By the construction in Definition 24, the embedded policy kernel of the refreshed insulin experiment is this matrix. Hence, for every pair of probability laws on , Lemma 1 gives Applying this with and is exactly Assumption 6 with .
The behavior probabilities are and the target policy is a probability law with . Therefore and Thus Assumption 5 holds with .
Let and be the selected stationary laws of and . The behavior and target policies in Definition 24 are probability vectors; together with the contraction estimate supplied by Lemma 1, Lemma 2 gives for and all . Write for the joint state at time . For and , define the nonrefresh action kernel as follows. If let be the probability that the Gaussian clipping-and-rounding rule in Definition 24 produces next glucose value , and set this probability to zero when is outside the glucose grid. Let with and equal to zero off their displayed supports. Conditional on the nonrefresh branch and action , Definition 24 draws the new diet and activity coordinates with laws and , shifts the old memory coordinates, and stores the current action. Hence the nonrefresh probability of moving from to under action is All factors in this display are nonnegative, so . The refresh branch in Definition 24 has probability and draws the next full state uniformly from the -point set , so its contribution to any fixed next state is The remaining branch has probability , and the policy-induced kernel averages the action-specific nonrefresh probabilities with weights . Thus, for every memoryless policy on , Since the behavior policy has nonnegative action probabilities, this identity gives the pointwise minorization The stationary-minorization calculation is then Since is a probability vector, Combining the last two displays yields and . This proves Assumption 7 with .
For the same horizon , write for the transition kernel of , and for its behavior and target policies, and for the stationary law induced by . The stationary immediate-weighting bias diagnostic is For this diagnostic on the refreshed finite insulin adaptation of Definition 24, Lemma 7 supplies the interval certificate which is the displayed bound
The five displayed conclusions are precisely the conjunction asserted for the refreshed insulin experiment.
∎References
- Hu, Yuchen and Wager, Stefan (2023). Off-Policy Evaluation in Partially Observed Markov Decision Processes under Sequential Ignorability. The Annals of Statistics. arXiv
- Mehrabi, Mohammad and Wager, Stefan (2025). Off-Policy Evaluation in Markov Decision Processes under Weak Distributional Overlap. . arXiv
- Zhang, Yuheng and Jiang, Nan (2024). On the Curses of Future and History in Future-dependent Value Functions for Off-policy Evaluation. Advances in Neural Information Processing Systems. arXiv
- Tennenholtz, Guy and Shalit, Uri and Mannor, Shie (2020). Off-Policy Evaluation in Partially Observable Environments. Proceedings of the AAAI Conference on Artificial Intelligence. doi
- Kallus, Nathan and Uehara, Masatoshi (2020). Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes. Journal of Machine Learning Research. arXiv
- Liao, Peng and Klasnja, Predrag and Murphy, Susan A. (2021). Off-Policy Estimation of Long-Term Average Outcomes with Applications to Mobile Health. Journal of the American Statistical Association. doi
- Luckett, Daniel J. and Laber, Eric B. and Kahkoska, Anna R. and Maahs, David M. and Mayer-Davis, Elizabeth J. and Kosorok, Michael R. (2020). Estimating Dynamic Treatment Regimes in Mobile Health Using {V}-Learning. Journal of the American Statistical Association. doi
- Maahs, David M. and Mayer-Davis, Elizabeth J. and Bishop, Freya K. and Wang, Lin and Mangan, Margo and McMurray, Robert G. (2012). Outpatient Assessment of Determinants of Glucose Excursions in Adolescents with Type 1 Diabetes: Proof of Concept. Diabetes Technology \& Therapeutics. doi
- Robins, James M. (1986). A New Approach to Causal Inference in Mortality Studies with a Sustained Exposure Period---Application to Control of the Healthy Worker Survivor Effect. Mathematical Modelling. doi
- Rosenbaum, Paul R. and Rubin, Donald B. (1983). The Central Role of the Propensity Score in Observational Studies for Causal Effects. Biometrika. doi
- Robins, James M. and Hern{\'a}n, Miguel A. and Brumback, Babette (2000). Marginal Structural Models and Causal Inference in Epidemiology. Epidemiology. doi
- Murphy, Susan A. (2003). Optimal Dynamic Treatment Regimes. Journal of the Royal Statistical Society: Series B (Statistical Methodology). doi
- Smallwood, Richard D. and Sondik, Edward J. (1973). The Optimal Control of Partially Observable Markov Processes over a Finite Horizon. Operations Research. doi
- Kaelbling, Leslie Pack and Littman, Michael L. and Cassandra, Anthony R. (1998). Planning and Acting in Partially Observable Stochastic Domains. Artificial Intelligence. doi
- Sutton, Richard S. and Barto, Andrew G. (2018). Reinforcement Learning: An Introduction. MIT Press.
- Precup, Doina and Sutton, Richard S. and Singh, Satinder P. (2000). Eligibility Traces for Off-Policy Policy Evaluation. Proceedings of the Seventeenth International Conference on Machine Learning.
- Jiang, Nan and Li, Lihong (2016). Doubly Robust Off-policy Value Evaluation for Reinforcement Learning. Proceedings of the 33rd International Conference on Machine Learning. arXiv
- Thomas, Philip S. and Brunskill, Emma (2016). Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning. Proceedings of the 33rd International Conference on Machine Learning. arXiv
- Liu, Qiang and Li, Lihong and Tang, Ziyang and Zhou, Dengyong (2018). Breaking the Curse of Horizon: Infinite-Horizon Off-Policy Estimation. Advances in Neural Information Processing Systems. arXiv
- Nachum, Ofir and Chow, Yinlam and Dai, Bo and Li, Lihong (2019). {DualDICE}: Behavior-Agnostic Estimation of Discounted Stationary Distribution Corrections. Advances in Neural Information Processing Systems. arXiv
- Uehara, Masatoshi and Huang, Jiawei and Jiang, Nan (2020). Minimax Weight and {Q}-Function Learning for Off-Policy Evaluation. Proceedings of the 37th International Conference on Machine Learning. arXiv
- Bennett, Andrew and Kallus, Nathan (2021). Proximal Reinforcement Learning: Efficient Off-Policy Evaluation in Partially Observed Markov Decision Processes. . arXiv
- Shi, Chengchun and Uehara, Masatoshi and Huang, Jiawei and Jiang, Nan (2022). A Minimax Learning Approach to Off-Policy Evaluation in Confounded Partially Observable Markov Decision Processes. Proceedings of the 39th International Conference on Machine Learning. arXiv
- Miao, Rui and Qi, Zhengling and Zhang, Xiaoke (2022). Off-Policy Evaluation for Episodic Partially Observable Markov Decision Processes under Non-Parametric Models. Advances in Neural Information Processing Systems. arXiv
- Uehara, Masatoshi and Kiyohara, Haruka and Bennett, Andrew and Chernozhukov, Victor and Jiang, Nan and Kallus, Nathan and Shi, Chengchun and Sun, Wen (2023). Future-Dependent Value-Based Off-Policy Evaluation in {POMDPs}. Advances in Neural Information Processing Systems. arXiv
- Zhang, Yuheng and Jiang, Nan (2025). Statistical Tractability of Off-Policy Evaluation of History-Dependent Policies in {POMDPs}. International Conference on Learning Representations. arXiv
- Kuang, Qi and Wang, Jiayi and Zhou, Fan and Qi, Zhengling (2025). Breaking the Order Barrier: Off-Policy Evaluation for Confounded {POMDPs}. Advances in Neural Information Processing Systems.
- Liao, Peng and Qi, Zhengling and Wan, Runzhe and Klasnja, Predrag and Murphy, Susan A. (2022). Batch Policy Learning in Average Reward Markov Decision Processes. The Annals of Statistics. doi
- Nahum-Shani, Inbal and Smith, Shawna N. and Spring, Bonnie J. and Collins, Linda M. and Witkiewitz, Katie and Tewari, Ambuj and Murphy, Susan A. (2018). Just-in-Time Adaptive Interventions ({JITAIs}) in Mobile Health: Key Components and Design Principles for Ongoing Health Behavior Support. Annals of Behavioral Medicine. doi
Comments on earlier versions
Anchored to: