theorem its proof invokes invoked by

CausalSmith · AI Causal Scientist — pid_slate_benefit_partialtransport_v1 · Partial ID · AI reviewer score 6.3/10 · pinned commit 8f179a2 · PDF · Lean code · Slides · arXiv · GitHub

Sharp Ordinal Benefit Bounds for Survivor Compliers under Selection and Noncompliance

Abstract

This paper characterizes sharp bounds on the survivor-complier strict-benefit probability in a finite ordered-outcome instrumental-variables model with treatment-induced selection. The target is the probability , the survivor-complier benefit probability, that treatment raises the ordered outcome among units whose treatment receipt is shifted by the instrument and whose outcome would be observed under either treatment state. Under conditional instrument independence, consistency and exclusion through received treatment, instrument overlap, treatment monotonicity, covariate-specific weak selection monotonicity, and positive aggregate survivor-complier mass, observed instrument contrasts identify two cellwise selected-complier outcome capacities. Their smaller total is the survivor-complier mass, including cells with zero selected-complier capacity gap. The sharp identified set is the interval obtained by solving a branch-free exact-mass partial-transport problem in each covariate cell and aggregating the resulting threshold-cut values. The lower and upper endpoints have closed finite formulas, and endpoint-attaining full latent laws establish sharpness. The same construction gives sparse endpoint witnesses computable in operations for finite covariate support and ordered outcome levels, with work when full coupling matrices are materialized. A three-level synthetic witness yields the sharp interval . For i.i.d. samples from the fixed finite-slate structural class with fixed finite support, prescribed screening thresholds satisfying and , a common instrument-overlap bound, and a uniform positive lower bound on aggregate survivor-complier mass, a projected screened plug-in estimator has a fixed-law Hadamard directional limit and a deterministic guarded confidence interval contains the full sharp identified interval uniformly.

Introduction

Many empirical questions concern a treatment’s probability of making a unit better off on an ordered scale. In a training-program setting, for example, a randomized offer can shift actual program receipt, employment determines whether wages are observed, and the outcome of interest may be a prespecified ordered wage category. A natural principal-stratum target is then the probability that program receipt raises the wage category among offer compliers who would be employed under either receipt state. This paper gives sharp bounds and guarded inference for that same-unit ordinal benefit probability in a finite ordered-outcome instrumental-variables model with treatment-induced selection.

The model combines a binary instrument , treatment receipt , selection , covariates , and an ordered outcome observed among selected units. The complier event is the group with , where are potential treatment receipts under the two instrument states. The survivor-complier target conditions further on , where are potential selection indicators under the two treatment states. The estimand is the probability that the treatment-one potential ordered outcome exceeds the treatment-zero potential ordered outcome among survivor compliers. In the Job Corps-style interpretation, is offer assignment, is program receipt, is employment, and is an ordered wage category; the numerical example below is a synthetic compatible law.

The identifying restrictions follow the local-average-treatment-effect and principal-stratification traditions (Imbens et al., 1994; Angrist et al., 1996; Frangakis et al., 2002). Conditional on covariates, the instrument is independent of potential treatments, potential selections, and potential ordered outcomes. Observed treatment equals treatment under the realized instrument, observed selection equals selection under the realized treatment, and observed outcomes equal potential outcomes under the realized treatment on selected units. Treatment receipt is monotone in the instrument, and instrument propensities satisfy cellwise overlap. Selection is governed by a covariate-specific weak monotonicity direction: within each complier cell, either treatment weakly raises selection or treatment weakly lowers selection. These conditions match the selected-complier logic used in IV sample-selection work (Chen et al., 2015; Kennedy et al., 2019; Dong et al., 2026).

The first result identifies the observable capacities that enter the benefit problem. For each covariate cell , the lower-arm capacities and upper-arm capacities are observed selected-complier outcome submeasures for and , respectively. Their totals and can differ because treatment changes selection. Weak selection monotonicity makes the survivor-complier mass in cell . Thus the observed law supplies two possibly unequal nonnegative capacities and an exact amount of mass that must be paired between them.

The construction follows a four-step route. Observed conditional instrument contrasts first produce the lower- and upper-arm selected-complier capacities. Their totals determine the cellwise exact survivor-complier mass. The threshold cuts then solve the exact-mass transport problem in each cell. Finally, the cellwise cut values are averaged with covariate-cell weights and normalized by aggregate survivor-complier mass to obtain the sharp interval.

The sharp-bound problem is a finite exact-mass partial-transport problem. In each cell, a feasible coupling places mass on ordered pairs , respects the row capacities , respects the column capacities , and has total mass . The cellwise benefit functional counts strict ordinal improvement. The paper proves that the feasible projections of full latent IV-selection laws are exactly these exact-mass transport polytopes, and that the sharp lower and upper values are the threshold-cut quantities and . Aggregating by covariate-cell probabilities gives Endpoint-attaining full latent laws show that the interval is sharp under the maintained structural law class.

The threshold formulas connect this construction to existing ordinal benefit bounds. Makarov-type and transport bounds characterize feasible joint distributions from marginal information (Makarov, 1982; Williamson et al., 1990; Fan et al., 2010; Villani, 2009). For complete ordered marginals, Lu et al. (2018) give closed-form sharp bounds for the strict-benefit probability. The present exact-mass formulation recovers that complete-marginal formula on the no-selection face where survivor-complier status coincides with complier status and the two capacity totals agree. Away from that face, the same threshold logic applies to selected-complier submeasures generated by noncompliance and treatment-induced selection. Related ordinal-benefit and probability-of-causation analyses study graphical, interventional, or unconfoundedness-based restrictions (Tian et al., 2000; Sachs et al., 2023; Duarte et al., 2023; Gabriel et al., 2024; de Aguas et al., 2025; Chen et al., 2026); the contribution here is the sharp survivor-complier IV-selection interval with endpoint-attaining latent laws.

The constructive proof is also an algorithm. The strict-benefit graph and its complement have nested threshold structure on a finite ordered support. A sparse threshold-flow routine computes all cellwise endpoint values and positive allocation traces in operations. Writing full lower and upper coupling matrices adds the expected materialization cost. A three-level witness makes the calculation explicit: one covariate cell, three ordered outcomes, full compliance, and treatment-increased selection produce capacities and , exact survivor-complier mass , and the sharp interval .

The estimation and inference results treat the sharp interval as a finite-dimensional functional of the observed-data law. The estimator replaces observed probabilities with empirical analogues, projects raw capacities onto the nonnegative cone, screens cells by empirical survivor-complier mass, and evaluates the same threshold formulas. At a fixed observed law, under independent sampling, overlap, the structural restrictions, positive aggregate survivor-complier mass, and a screening threshold with and , the retained support is recovered with probability tending to one, the plug-in endpoints are consistent, and the scaled endpoint error converges to the Hadamard directional derivative of the reduced-support endpoint map applied to the centered multinomial Gaussian limit. This accommodates simultaneous active threshold cuts and zero selected-complier capacity gaps through the directional derivative framework for nonsmooth maps (van der Vaart, 1998; Fang et al., 2019; Hong et al., 2018).

For finite-sample containment, the paper gives a deterministic guard. Over a finite event collection determined by and , a union-bound deviation threshold yields an explicit capacity-and-screening envelope. Under fixed instrument overlap and a uniform lower bound on aggregate survivor-complier mass, padding the screened plug-in endpoints by this envelope produces a guarded interval that contains the entire sharp identified interval with probability at least , uniformly over the indexed finite-slate law family and every sample size. The guard radius is , and the choice with and gives radius . The technical appendix records supporting derivations and cited inputs.

The paper proceeds as follows. Section 2 discusses the related econometric literature. Section 3 introduces the potential-outcome setup, the structural restrictions, and the observable capacity identities. Section 4 states the exact-mass partial-transport representation, the threshold-cut endpoint formulas, the endpoint-attaining law construction, and the no-selection reduction. Section 5 gives the sparse threshold-flow algorithm and the three-level witness. Section 6 develops the projected screened estimator, the fixed-law directional limit, and the deterministic guarded confidence interval. Section 7 records limitations and future work, and the technical appendix collects supporting derivations.

Related work

This paper works in a potential-outcome setting with a binary instrument, treatment-induced selection, and a finite ordered outcome. Its target combines three strands of econometric work: instrumental variables and principal strata, sample selection bounds, and sharp bounds for ordinal benefit probabilities.

The instrumental-variables component follows the local-treatment-effect tradition in which a binary instrument shifts treatment take-up and identifies causal information for compliers (Imbens et al., 1994; Angrist et al., 1996). Principal stratification gives the language for conditioning on joint potential statuses, including compliance and post-treatment selection strata (Frangakis et al., 2002). Partial-identification results for IV models, including sharp bounds under latent response restrictions, provide the surrounding identification perspective (Balke et al., 1997; Manski, 1990; Manski, 1997; Manski et al., 2000; Manski, 2003). The present analysis uses that tradition to study survivor compliers, a principal stratum defined jointly by compliance behavior and selection under treatment states, and asks for sharp bounds on the probability that treatment strictly improves an ordered outcome among that group.

Selection is central because the ordered outcome is observed on selected units. Classical sample-selection models motivate the econometric concern that outcome comparisons among selected units mix treatment effects with selection effects (Heckman, 1979). Lee-style trimming bounds give a transparent solution under monotone selection (Lee, 2009), and subsequent work develops this approach for instrumental variables and principal strata in settings with attrition or post-treatment selection (Chen et al., 2015; Blanco et al., 2013; Chen et al., 2018; Kennedy et al., 2019). The closest selection-bounds comparison is Dong et al. (2026), who study covariate-specific weak selection monotonicity and selected-complier outcome distributions. Closest on the benefit side under selection is Possebom et al. (2025), who bound the probability of causation for a binary outcome when selection determines whether the outcome is recorded, and who identify their target on the always-observed stratum. The present paper differs on both axes that generate its mathematics: the outcome is ordered with levels, so the target is a same-unit comparison whose sharp value is a coupling problem rather than a pair of marginal probabilities; and the conditioning stratum is the survivor compliers, whose two arms carry unequal capacity totals when treatment shifts selection. That inequality is what forces the exact-mass partial-transport formulation below: the two arms must be coupled subject to a common exact-mass constraint, so the sharp value solves a transport problem rather than following from a pair of marginal probabilities. The present paper uses the same type of covariate-specific selection direction as an input to a sharp ordinal-benefit analysis for survivor compliers.

The ordinal outcome component connects to work on distribution-free bounds for functions of paired potential outcomes. Makarov-type bounds characterize feasible joint distributions given marginal distributions (Makarov, 1982; Williamson et al., 1990), and optimal-transport formulations give a useful finite-dimensional view of the same coupling problem (Villani, 2009; Fan et al., 2010). For ordered categorical outcomes, the benefit probability is a functional of the unobserved joint distribution of two potential ordered outcomes; complete marginal information yields closed-form formulas for this functional (Lu et al., 2018; Lu et al., 2020). This paper adapts that coupling logic to the selected-complier capacities generated by an IV model with treatment-induced selection: within each covariate cell, the relevant margins are selected-complier subprobability capacities with potentially unequal totals.

Recent work also develops sharp ordinal benefit bounds under other observable restrictions. Gabriel et al. (2024) study sharp ordinal-benefit bounds including response-function IV restrictions, and de Aguas et al. (2025) establish sharp bounds for tiered benefit under monotonicity with inference. Related causal-graphical analyses of probability-of-causation and benefit-type estimands include Tian et al. (2000); Sachs et al. (2023); Duarte et al. (2023). Chen et al. (2026) study principal generalized causal effects under treatment unconfoundedness and ignorability conditions. The contribution here is a sharp identified interval for a survivor-complier strict-benefit probability under the paper’s IV, consistency, overlap, no-defiers, weak-selection-monotonicity, and positive survivor-mass conditions, together with a finite-dimensional selected-complier capacity and partial-transport representation used for computation and inference.

The inference part of the paper belongs to the econometric literature on partially identified parameters and directionally nonregular maps. Confidence procedures for identified sets motivate coverage statements for an interval-valued target (Imbens et al., 2004; Chernozhukov et al., 2007; Romano et al., 2010; Tamer, 2010; Molinari, 2020; Chernozhukov et al., 2013). Directional differentiability and bootstrap validity for nonsmooth functionals provide the local asymptotic framework (van der Vaart, 1998; Fang et al., 2019; Hong et al., 2018). Recent contributions on inference for bounds and partially identified causal parameters, including Russell (2021), Levis et al. (2025), Ben-Michael (2025), Khandamiryan et al. (2026), Noack (2026), and Lee et al. (2024), frame the inferential goal: to report uncertainty for a set identified by observable distributions and maintained structural restrictions. Here that goal is implemented through a guarded finite-sample construction tailored to the threshold-cut representation of the survivor-complier benefit interval.

Setup and assumptions

Throughout the paper, covariate support is finite and sums over covariate cells range over . Outcome levels are ordered as , and indices range over this ordered support unless stated otherwise. Generic constants are positive and finite, Euclidean norms are used for finite-dimensional vectors, and asymptotic notation is always along the sampling sequence introduced in the inference section.

We begin with the potential-outcome structure. The observed system has covariates , a binary instrument , treatment receipt , selection , and an ordered outcome observed on selected units. The model also fixes the latent probability law , the finite covariate carrier , the outcome-support size , and the ordered outcome set .

The section proceeds from variables to restrictions to capacities: the slate fixes the variables, the consistency and exclusion restrictions connect observed cells to potential outcomes, monotonicity identifies the complier and survivor-complier components, and the observable-capacity definition packages the two selected-complier submeasures used by the transport problem.

Definition 1 [synth_13] (POSlate system structure).

A POSlate system over a potential-outcome system , a measurable covariate carrier , and an outcome-support size consists of:

  • (Latent space.) The latent probability space of is standard Borel.

  • (Outcome support.) The integer satisfies , and the ordered outcome support is .

  • (Selected nodes.) Five pairwise distinct nodes of , written .

  • (Value spaces.) The node is measurably identified with , while , , and are measurably identified with .

  • (Outcome node.) The node is measurably identified with the finite ordered support .

Here is the binary instrument, is the covariate, is treatment receipt, is selection, and is the finite ordered outcome.

⊢ Lean

Potential treatments, selections, and outcomes are indexed by the intervention that defines them. The treatment potentials vary the instrument, while the selection potentials and outcome potentials vary treatment receipt.

Definition 2 [synth_11] (Control potential outcome).

The symbol , the potential ordered outcome under treatment zero, is the map from the sample space to the ordered outcome levels obtained by setting treatment to zero:

⊢ Lean
Definition 3 [synth_10] (Selection under treatment one).

The potential selection indicator under treatment one is the map defined by

⊢ Lean
Definition 4 [synth_9] (Treatment-zero selection).

Define the potential selection indicator under treatment zero by Equivalently, is the selection variable evaluated counterfactually when the treatment variable is set to .

⊢ Lean
Definition 5 [synth_8] (Potential treatment under instrument one).

Let , the potential treatment under instrument state one, be the Boolean unit-level variable on given by the treatment value that would be realized under the intervention . Equivalently, for each unit ,

⊢ Lean
Definition 6 [synth_7] (Potential treatment under zero).

The potential treatment under instrument value zero is the binary variable defined by

⊢ Lean
Definition 7 [synth_12] (Treatment-one outcome).

The potential ordered outcome under treatment one is the map given by evaluating the ordered outcome counterfactual under the intervention that sets treatment receipt to .

⊢ Lean

The first group of identifying restrictions is the standard IV structure with treatment-mediated selection and outcomes. It separates instrument assignment from latent response types within covariate cells, and it connects the observed variables to their potential-outcome counterparts.

Assumption 1 [ass:iv-independence] (Instrument independence).

The potential treatments , potential selection indicators , and potential ordered outcomes are conditionally independent of the binary instrument given the covariate cell :

⊢ Lean

Assumption 1 is the conditional instrument independence condition: after conditioning on , variation in is as good as randomly assigned for the potential treatments, selections, and ordered outcomes (Angrist et al., 1996).

Assumption 2 [ass:treatment-consistency] (Treatment consistency).

Observed treatment receipt equals the potential treatment under the realized instrument state: almost surely.

⊢ Lean

Assumption 2 is the treatment consistency condition, which links realized receipt to the potential treatment indexed by the realized instrument (Angrist et al., 1996). Together with independence, it gives observed first-stage contrasts their usual complier interpretation.

Assumption 3 [ass:selection-exclusion-consistency] (Selection consistency).

Observed selection equals the potential selection indicator under the realized treatment: almost surely.

⊢ Lean

Assumption 3 is the selection exclusion through received treatment condition: instrument assignment affects selection through the treatment actually received (Kennedy et al., 2019).

Assumption 4 [ass:outcome-exclusion-consistency] (Outcome consistency).

Among selected units, the observed outcome equals the potential ordered outcome under the realized treatment: on almost surely.

⊢ Lean

Assumption 4 is the outcome exclusion through received treatment condition on selected units (Kennedy et al., 2019). Because outcomes are observed when , this condition fixes the outcome content of selected observed cells.

The observed-data law records the variables available to the analyst, with the outcome marked missing outside the selected sample. We write the observed tuple as and its law as .

Definition 8 [synth_17] (Observed finite data law).

The observed-data map sends each latent unit to the finite observed cell The observed-data law is the pushforward of the latent law through this observed-data map:

⊢ Lean

Cell weights enter every aggregate target below. The covariate-cell probability is defined before the complier and survivor-complier quantities are formed.

Definition 9 [synth_14] (Covariate cell mass).

For each covariate cell , the cell probability is

⊢ Lean
Definition 10 [synth_1] (Complier event definition).

The complier event is the set of units whose potential treatment equals when and equals when :

⊢ Lean

The complier event is the latent stratum whose treatment receipt is shifted by the instrument. The benefit parameter studied later is restricted further to compliers who would be selected under both treatment states.

For compact reference, the following definition collects the finite slate, the observed law, the monotonicity structure, and the cellwise capacity objects used throughout the rest of the paper.

Definition 11 [synth_22] (Slate-benefit partial-transport system).

A slate-benefit partial-transport system is a POSlate potential-outcome structure with finite discrete covariate support , ordered outcome support for , binary instrument , binary treatment , selection indicator , potential treatments , potential selections , and potential outcomes . The observed variables satisfy , , and on , with observed element and observed law . The system carries a latent law under which , almost surely, and for each cell with , where , a direction satisfies . Its cellwise observable capacities are and , with , aggregate survivor-complier mass , and survivor-complier outcome subcouplings constrained by the exact-mass partial-transport sets .

The remaining assumptions impose overlap and monotonicity. The instrument propensity and the overlap constant determine where conditional instrument contrasts are well defined.

Definition 12 [synth_16] (Instrument propensity by cell).

For each covariate cell , the conditional instrument propensity is the real-valued conditional probability of in that cell:

⊢ Lean
Assumption 5 [ass:instrument-overlap] (Instrument propensity overlap).

For a slate-benefit partial-transport system, the instrument-overlap condition at the instrument-overlap constant consists of:

  • (Positivity.) .

  • (Half-overlap.) .

  • (Cellwise overlap.) For every covariate cell with covariate-cell probability , the conditional instrument propensity satisfies

⊢ Lean

Assumption 5 is the conditional instrument overlap condition (Kennedy et al., 2019). It keeps both instrument states present in every supported covariate cell, so the observed conditional contrasts used to form capacities are defined cell by cell.

Assumption 6 [ass:no-defiers] (Treatment monotonicity).

Potential treatment receipt is monotone in the instrument: almost surely.

⊢ Lean

Assumption 6 is the instrument monotonicity condition (Angrist et al., 1996). It orients the first stage so that the shifted latent group is the complier stratum in Definition 10.

Assumption 7 [ass:weak-selection-monotonicity] (Weak selection monotonicity).

For every covariate cell with , the covariate-specific direction satisfies

⊢ Lean

Assumption 7 is the covariate-specific weak sample-selection monotonicity condition of Dong et al. (2026). The direction is allowed to vary across covariate cells, and the condition makes the sign of the selected-complier selection contrast interpretable within cells that contain compliers.

Assumption 8 [ass:positive-aggregate-survivors] (Positive aggregate survivors).

The aggregate survivor-complier mass satisfies

⊢ Lean

Assumption 8 is the positive target-stratum probability condition (Kennedy et al., 2019). It ensures that the aggregate survivor-complier mass can normalize the benefit probability studied in the next section.

The maintained structural law class is the collection of latent laws satisfying the restrictions just stated.

Definition 13 [def:structural-law-class] (Structural law class ).

The structural law class is This is the finite-ordered, finite-covariate specialization of the corresponding Dong–Heiler cellwise restrictions on cells with . Cells with have zero capacities and are inert for the target.

⊢ Lean

The observable selected-complier capacities are the bridge from the observed law to the partial-transport problem. For each cell, records lower-arm selected-complier mass at outcome level , while records upper-arm selected-complier mass at level . Their totals and , their difference , and their exact survivor-complier mass determine the feasible cellwise coupling sets introduced in the next section.

Definition 14 [def:observable-capacities] (Observable capacities ).

For each and , let and be the lower-arm and upper-arm selected-complier outcome capacities obtained from the corresponding conditional instrument contrasts. Define the capacity totals, the capacity-total gap, and the exact survivor-complier mass by

⊢ Lean

The section closes with the identification statement for these capacities. It converts the observed IV contrasts into latent selected-complier submeasures and identifies the exact survivor-complier mass cell by cell.

Theorem 1 [prop:capacity-identification] (Capacity identification).

Let be a five-node potential-outcome slate with covariate , binary , ordered outcome support with , and standard-Borel latent probability space. Let be the instrument-overlap constant, and let be the covariate-specific weak selection-monotonicity direction. Suppose that:

  • (IV independence.) Assumption 1 holds.

  • (Treatment consistency.) Assumption 2 holds.

  • (Selection exclusion and consistency.) Assumption 3 holds.

  • (Outcome exclusion and consistency.) Assumption 4 holds.

  • (Instrument overlap.) Assumption 5 holds with constant .

  • (No defiers.) Assumption 6 holds.

  • (Weak selection monotonicity.) Assumption 7 holds with direction .

Let , , , and be the observable capacity contrasts formed from the observed-data law. Then the observable capacities are nonnegative in every covariate cell, all covariate-cell probabilities satisfy , and, for every supported cell with , for every , for every , and Moreover, for every supported , implies , implies , and makes both directions observationally admissible under the observed-data law.

⊢ Lean

The intuition is that the IV independence and consistency restrictions express selected outcome contrasts by latent complier components, while treatment monotonicity fixes which treatment response type contributes to the contrast. Weak selection monotonicity then aligns the selection contrast so that the smaller selected-complier total is exactly the always-selected complier mass in that cell. Consequently, Theorem 1 supplies the two nonnegative marginal submeasures that feed the exact-mass partial-transport construction below.

Sharp benefit bounds

The capacity identification result in Theorem 1 reduces the target parameter to a finite collection of cellwise coupling problems. In each covariate cell, the lower-arm and upper-arm selected-complier outcome capacities provide two submarginals, while the exact survivor-complier mass fixes how much of these capacities must be paired. The resulting object is the exact-mass partial-transport polytope , whose elements are cellwise survivor-complier coupling arrays . This is the finite ordinal analogue of a partial transport problem with capacity constraints (Villani, 2009).

Definition 15 [def:partial-transport-polytope] (Partial-transport polytopes ).

For each , define The comparison polytopes are and The cellwise survivor-complier coupling polytope is

⊢ Lean

The comparison polytopes in Definition 15 make explicit which side of the two selected-complier capacity totals binds. When the upper-arm total exceeds the lower-arm total, the row constraints can bind; when the lower-arm total exceeds the upper-arm total, the column constraints can bind. The branch-free notation records that the paper’s identified coupling set is always the exact-mass set, with the operative branch determined directly by the observed capacity totals.

The endpoint calculations use ordinal threshold cuts. The lower endpoint is obtained by forcing enough mass onto weakly adverse pairs to limit strict benefit, while the upper endpoint packs as much mass as possible onto pairs with treatment-one outcome above treatment-zero outcome. The finite support lets these constraints be summarized by prefixes and tails, in the spirit of sharp distributional bounds for ordinal and discrete outcomes (Makarov, 1982; Williamson et al., 1990; Fan et al., 2010).

Definition 16 [def:threshold-cuts] (Threshold cuts ).

For every , define the prefix and tail sums The lower and upper threshold-cut values are and These formulas apply on positive-survivor zero-gap cells and on zero-survivor cells.

⊢ Lean

The lower threshold-cut value and upper threshold-cut value are unnormalized benefit-mass bounds within a cell. Their formulas combine the ordinal cut constraints with the exact survivor-complier mass, including zero-gap cells where the two capacity totals coincide.

Definition 17 [def:identified-interval] (Identified interval ).

The endpoint map is The sharp identified interval for the benefit probability is

⊢ Lean

Definition 17 aggregates the cellwise cut values into the endpoint map . The lower endpoint , upper endpoint , and interval are normalized by the aggregate survivor-complier mass from Assumption 8, so the estimand is the fraction benefiting among survivor compliers.

Definition 18 [synth_4] (Threshold-cut capacity aggregates).

For each covariate cell and threshold , define the lower-arm prefix, strict-prefix, and upper-tail capacities by , , and . Define the corresponding higher-arm prefix and upper-tail capacities by and .

The aggregate notation in Definition 18 records the prefix and tail capacities used by the threshold formulas. The additional upper tail , together with , , , and , gives the algebraic identities needed to compare the exact-mass problem with complete-marginal ordinal bounds.

Theorem 2 [prop:tie-face-collapse] (Tie face collapse).

Let be a five-node potential-outcome system as in Definition 13, with ordered outcome support of size . Fix an instrument-overlap constant and a covariate-specific weak selection-monotonicity direction . Assume:

  • (Structural restrictions.) The law satisfies Assumptions 1, 2, 3, 4, 5, 6, 7, and 8, with instrument-overlap constant and weak selection-monotonicity direction .

  • (Observable capacities.) Let be the observable capacity vector formed from the observed-data law as in Definition 14, and let its validity be the validity delivered by Theorem 1.

Then, for every covariate cell with , For every such cell , if , then Under the same zero-gap condition, Finally, for every such cell , if , then for every threshold ,

⊢ Lean

Theorem 2 characterizes the zero-gap face of the partial-transport problem. At , the exact survivor-complier mass equals both selected-complier capacity totals, so the row-exact, column-exact, and doubly exact comparison polytopes coincide with the maintained coupling polytope. The threshold identity then expresses the same balance through prefixes and tails, which lets implementations evaluate zero-gap cells using the branch-free formulas in Definition 16.

The next construction turns the threshold formulas into full latent laws. Cellwise, it first builds extremal couplings for the strict-benefit and nonbenefit graphs; globally, it embeds those couplings into potential-outcome laws that retain the observed law and the structural restrictions.

Algorithm 1 [def:threshold-flow-construction] (Threshold-flow construction ).

The inputs are the observable cellwise capacities and , the capacity totals and , and the threshold-cut values and .

  1. For each , compute the exact mass and the threshold sums:

  2. Compute the threshold-cut values and

  3. For the upper endpoint, construct a maximum subflow on the nested benefit graph . Extend it to total mass by allocating the residual mass on the complete bipartite graph.

  4. For the lower endpoint, construct a maximum subflow on the nested nonbenefit graph . Extend it to total mass by allocating the residual mass on the complete bipartite graph.

  5. The resulting coupling in each cell satisfies

  6. Allocate row or column residual mass to the one-sided selected-complier stratum determined by the sign of Place the remaining complier mass in the never-selected stratum and adjoin the unchanged never-taker and always-taker components of a baseline compatible law.

  7. The construction outputs endpoint-attaining latent laws and with endpoint map

⊢ Lean

The lower endpoint law and upper endpoint law are constructed cell by cell and then completed with the compliance and selection strata. The residual-mass step aligns the partial-transport coupling with the weak selection-monotonicity direction from Assumption 7, while the unchanged never-taker and always-taker components preserve the observed IV law.

Theorem 3 [thm:full-law-endpoint-attainment] (Endpoint law attainment).

Let be the observed-data law induced by a five-node potential-outcome slate system with covariates , binary instrument , treatment , selection , and ordered outcome , where . Let encode the covariate-specific weak selection-monotonicity direction, with corresponding to and corresponding to . Let be the instrument-overlap constant. Suppose that:

  • (Structural restrictions.) The latent slate system satisfies Assumption 1, Assumption 6, and Assumption 7.

  • (Tie-safe survivor model.) The same law satisfies Assumption 2, Assumption 3, Assumption 4, Assumption 5 at , Assumption 6, and Assumption 7 in direction .

  • (Positive aggregate mass.) For the observable capacity vector formed from as in Definition 14, the cell probabilities , and the survivor-complier masses ,

Then there exists a baseline law compatible with and : it induces , has observable capacities , has conditional complier mass in every cell at least , and carries the same never-taker and always-taker components used by the latent completion. For this baseline, the cellwise threshold-flow construction in Algorithm 1 returns two full laws and over . These laws are endpoint witnesses for the endpoint map : each induces exactly , satisfies the structural restrictions with instrument-overlap constant and direction , and attains Moreover, for every covariate cell , realizes the lower threshold latent completion generated from the baseline at , and realizes the upper threshold latent completion generated from the baseline at .

⊢ Lean

Theorem 3 establishes that the cellwise extremal couplings are attainable by full potential-outcome laws. The construction first preserves the observed capacities identified in Theorem 1, then assigns the unmatched selected-complier capacity to the one-sided selected stratum allowed by weak selection monotonicity. Because the theorem attains the two endpoint probabilities under the maintained structural restrictions, the threshold formulas are realized by feasible full latent laws.

The exact-mass formulation also contains the familiar complete-marginal ordinal benefit problem on the no-selection face. In that submodel, survivor-complier status coincides with complier status in every relevant cell, so the two selected-complier capacity totals agree and the partial-transport polytope becomes a doubly exact coupling set.

Theorem 4 [thm:no-selection-reduction] (No-selection reduction).

Let the ordered outcome support have levels, and let the observed capacities be formed from the observed-data law as in Theorem 1. Suppose that:

  • (Structural law.) The latent law satisfies the structural restrictions in Definition 13, including Assumption 1 and Assumption 6, with instrument-overlap constant and weak selection-monotonicity direction .

  • (No selection submodel.) For every covariate cell with ,

Then, for every covariate cell with , Moreover, if , then the normalized threshold-cut endpoints satisfy and

⊢ Lean

The normalized formulas in Theorem 4 match the closed-form strict-benefit bounds for two complete ordinal marginals in Lu et al. (2018). The theorem identifies the precise face of the present model where that comparison applies: every relevant complier is always selected, the capacity gap is zero, and the exact-mass polytope has both margins fixed.

It remains to name the cellwise objective optimized over the coupling set.

Definition 19 [synth_5] (Cellwise benefit-mass functional).

For a survivor-complier outcome subcoupling in cell , define the unnormalized benefit mass by . Thus counts exactly the mass assigned to outcome pairs whose treatment-one outcome level strictly exceeds the treatment-zero outcome level.

The functional is linear in the coupling and counts strict ordinal improvement. With this objective, the threshold cuts in Definition 16 become the exact lower and upper values of the cellwise optimization problem.

Theorem 5 [thm:sharp-exact-mass-threshold-interval] (Sharp exact-mass interval).

Let be the observed-data law generated by a five-node potential-outcome slate with ordered outcome support of size . Suppose the latent law satisfies the structural restrictions in Definition 13 for the instrument-overlap constant and weak-selection direction . Let be the observable capacity vector from Definition 14, identified from as in Theorem 1, and suppose that

Then the following statements hold.

  • (Sharp cellwise projection.) For every cell with , the set of survivor-complier couplings generated by feasible full laws with observed law is exactly

  • (Exact-mass branch.) For every cell with ,

  • (Threshold endpoints.) For every cell with , Writing and for the lower and upper cellwise threshold-flow couplings from Algorithm 1,

  • (Zero-mass cells.) For every cell with ,

Consequently, the sharp identified interval for the benefit probability over feasible full laws is

⊢ Lean

Theorem 5 is the section’s sharpness result. It characterizes the projection of feasible full laws onto the exact-mass partial-transport polytope, identifies the sign-specific comparison polytope selected by the capacity gap, and evaluates the lower and upper values of the strict-benefit objective by the threshold cuts. Together with endpoint law attainment in Theorem 3, these statements deliver the sharp identified interval for the survivor-complier benefit probability under the structural law class in Definition 13.

Computation and witness

The threshold-cut formulas in Definition 16 and the endpoint construction in Algorithm 1 have two practical consequences. First, the construction can be implemented cell by cell by sorting the ordered outcome levels once and moving through nested lower and upper threshold sets. Second, the same calculation produces explicit sparse couplings, so the numerical endpoint is accompanied by a witness allocation in the partial-transport polytope.

We begin with a three-level law that makes the objects in Definition 14 concrete. The example has one covariate cell, three ordered outcomes, and full compliance. Selection differs across treatment states, so the lower- and upper-arm selected-complier capacities have unequal totals before the exact survivor-complier mass is imposed.

Definition 20 [def:three-level-witness] (Three-level witness and ).

The observed law and latent law are defined as follows. The covariate satisfies almost surely, the binary instrument satisfies and the ordered outcome has levels. Conditional on , with all other conditional cells assigned probability zero. Conditional on , with all other conditional cells assigned probability zero.

The latent law assigns every unit potential treatments and . It allocates masses to units with potential selection indicators and diagonal potential-outcome pairs respectively. It allocates mass to units with and , with residual masses on outcome levels . It allocates mass to units with . The instrument is independent of these potential variables with probability , and latent outcomes are assigned arbitrarily wherever selection makes them unobserved.

⊢ Lean

The observed witness law displays all observed cells needed to recover the capacity vector, while the latent witness law supplies a compatible potential-outcome system. The selected always-survivor mass is diagonal in the ordered outcome, and the additional selected mass under treatment one raises the upper-arm capacities. The resulting calculation is small enough to audit directly but still contains the capacity imbalance that drives a nondegenerate sharp interval.

Theorem 6 [prop:three-level-witness] (Three-level witness sharpness).

For the single-cell three-level witness, let be the observed-data law assigning mass to the eight displayed observed cells, in order. Let be the observable capacity vector obtained from . Then the following assertions hold:

  • (Full-law realization.) There exists a latent potential-outcome system with a three-level slate system such that the survivor model is tie-safe at instrument overlap with the single covariate cell retained, its observed law is , its encoded full law is the displayed full witness law, and almost surely, and almost surely.

  • (Observable capacities.) At the single covariate cell ,

  • (Derived totals and cuts.) The same capacity vector satisfies and

  • (Identified interval.) With the unit cell weight , the nonnegative-capacity validity certificate, positive aggregate mass, and the observed-law-domain certificate for , the sharp identified interval is

  • (Endpoint attainment.) There exist latent laws and , each with a three-level slate system, satisfying the endpoint-witness conditions for , overlap , and the retained single covariate cell, such that attains and attains .

⊢ Lean

Theorem 6 shows how the observed margins translate into the interval . In the witness, the exact survivor-complier mass is , while the upper threshold cut is . Dividing the upper cut by the aggregate survivor-complier mass gives , and the lower cut gives the lower endpoint . Endpoint-attaining latent laws therefore realize both ends of the interval within the same IV-selection model.

The computational routine implements the same threshold logic used in the sharpness theorem. It computes the two cut values, records the positive allocations that attain them, and separates the sparse work needed for endpoint evaluation from the additional work required to materialize full coupling matrices.

Definition 21 [synth_20] (Costed threshold-flow routines).

For finite , , and nonnegative capacity arrays and , the costed sparse threshold-flow routine is the cellwise algorithm that computes , , , , the prefix and tail sums, and the threshold values and ; constructs an upper-attaining subflow greedily on the nested graph and a lower-attaining subflow greedily on the nested graph ; and completes each subflow to total mass by a two-pointer allocation on the residual complete bipartite graph. Its output is together with sparse allocation traces listing exactly the positive row–column allocations produced for the lower- and upper-attaining couplings in . The actual sparse operation count is the number of arithmetic operations, comparisons, pointer advances, and trace writes executed by this sparse implementation. The charged sparse operation count assigns one charge to each prefix or tail update, threshold comparison, pointer advance, residual-capacity update, and positive-allocation trace append. The costed dense threshold-flow routine performs the same threshold-flow construction and then materializes each lower- and upper-attaining coupling matrix, writing zero entries as explicit matrix entries. Its actual dense operation count is the number of arithmetic operations, comparisons, pointer advances, residual-capacity updates, and matrix writes executed by the dense implementation. Its charged dense operation count is the sparse charge plus one charge for each materialized matrix entry.

Definition 22 [synth_21] (Sparse threshold-flow operation counts).

For a covariate cell , let be the number of arithmetic operations and comparisons actually performed by the costed sparse threshold-flow routine of Algorithm 1 in cell , including the computation of , the prefix and tail sums, the threshold values and , and the sparse lower- and upper-attaining couplings. Let be the operation count assigned by the routine’s charging scheme in cell , where each pass over an outcome level and each pointer advance or allocation step in the greedy nested-graph subflow and residual two-pointer completion is charged to the corresponding sparse routine step.

These definitions count the operations at the level used by the constructive proof of Theorem 5. The sparse trace is the economically relevant certificate: it lists the row–column masses that attain the endpoint. Dense matrices are useful when an implementation or diagnostic display requires a complete array, and their cost includes the explicit zero writes.

Theorem 7 [thm:linear-sparse-threshold-flow] (Sparse threshold flow).

There exist positive integer constants and such that the following holds. Let be a finite covariate support, let , and let and , the lower- and upper-arm capacity arrays, be indexed by and . Suppose for every and every . For the threshold-flow construction in Algorithm 1, write and for the lower and upper threshold-cut values in Definition 16. Then:

  • (Sparse operation bounds.) The actual sparse operation count and the charged sparse operation count are each bounded above by

  • (Dense operation bounds.) The actual dense operation count and the charged dense operation count are each bounded above by

  • (Magnitude-invariant charged schedules.) For any other nonnegative capacity arrays on the same and the same , the charged sparse operation counts are equal and the charged dense operation counts are equal.

  • (Cellwise output.) For every , the costed sparse threshold-flow routine returns the threshold-cut pair and the sparse lower and upper allocation traces from Algorithm 1. The induced lower and upper couplings and satisfy belong to the branch-free partial-transport polytope in Definition 15, and attain the threshold-cut benefit masses

⊢ Lean

Theorem 7 establishes that exact endpoint computation scales linearly in the number of outcome levels when the output is kept sparse, uniformly across covariate cells. The sparse bound follows the nested-threshold structure: each pass updates a prefix or tail quantity, and each pointer movement is charged once. A quadratic term enters precisely when the implementation writes full lower and upper matrices. The magnitude-invariance clause is useful for repeated evaluation because the charged schedule depends on the support size and outcome grid, while the returned cut values and allocation masses depend on the realized capacities.

Estimation and guarded inference

The sharp interval in Definition 17 is a finite-dimensional functional of the observed-data law. Estimation therefore proceeds by replacing the observable probabilities in Definition 14 with empirical analogues, projecting the resulting capacities onto their nonnegative range, and evaluating the same threshold-cut formula cell by cell. The inferential problem is the usual one for nonsmooth partially identified endpoints: max, min, screening, and quotient operations can all determine the local behavior of the endpoint map (Imbens et al., 2004; Chernozhukov et al., 2007; Tamer, 2010; Molinari, 2020; Fang et al., 2019). The construction below keeps these operations explicit.

The finite-sample guard uses two sampling conditions. The first keeps the target stratum uniformly away from zero in aggregate, so that the endpoint quotients remain stable. The second fixes the sampling scheme for the empirical observed law.

Assumption 9 [ass:uniform-aggregate-survivor-bound] (Uniform aggregate survivor bound).

The uniform aggregate lower bound is strictly positive and bounds the aggregate survivor-complier mass in Definition 17 from below:

⊢ Lean

The requirement in Assumption 9 is a uniform positive target-stratum probability condition: the aggregate survivor-complier mass entering the denominator of the identified endpoints is bounded below by . This is the standard positivity requirement for inference on survivor-type causal targets (Kennedy et al., 2019).

Assumption 10 [ass:iid-sampling] (I.i.d. sampling).

The observations are independent draws from the observed-data law .

⊢ Lean

Assumption 10 is the independent sampling condition. It supplies the empirical-process input for the finite observed support and is the conventional baseline for the delta-method and multiplier arguments used below (van der Vaart, 1998).

Definition 23 [def:uniform-law-class] (Uniform law class ).

The uniform law class at index is

⊢ Lean

The class Definition 23 collects the laws for which the structural restrictions in Definition 13, the lower aggregate-mass condition, and independent sampling hold at the sample-size index. The notation will be used for uniform probability statements indexed by .

Definition 24 [synth_15] (Empirical cell probability).

For observations taking values in observed-data records with covariate cell component , realization , sample size , and covariate cell , the empirical covariate-cell probability is

⊢ Lean

The empirical covariate-cell probability is the sample analogue of the cell weight in Definition 17. It is the outer weight in the plug-in sums and also enters the screening score.

Algorithm 2 [def:plugin-endpoint-estimator] (Plug-in endpoint estimator ).

Given observations , a finite covariate support , outcome support , and screening threshold , define the projected screened plug-in endpoint estimator by the following steps.

  1. For each empirical conditional probability given , use the empirical ratio when the empirical denominator is positive and use the fixed value zero when that denominator is zero. These empirical conditional probabilities form the raw capacity vector

  2. Define the nonnegative capacity cone and the projected empirical capacity vector by With the fixed Euclidean metric, the minimizer is unique, and is the coordinatewise positive-part map. Write

  3. For each , define the empirical capacity totals, capacity-total gap, and empirical exact mass by

  4. For each , define the empirical lower and upper threshold-cut values by and

  5. Define the empirical retained-cell mass score and retained-cell indicator by

  6. Define the screened empirical aggregate mass and screened empirical endpoint numerators by

  7. Output the projected screened plug-in endpoint estimator

⊢ Lean

Algorithm 2 mirrors the population construction in Definition 17. The raw empirical capacity vector is projected to the nonnegative capacity cone, giving the projected empirical capacity vector . The retained-cell score and screening threshold restrict the quotient to cells with empirical survivor-complier mass large enough to enter stably. The resulting estimator is the sample analogue of the endpoint map.

For diagnostics and finite-sample containment, the guard works with a finite collection of observable events and a multiplier process over observed cells. The operational event collection consists of cell events , arm events , and selected-outcome events , indexed by , , and . Hence

Definition 25 [synth_18] (Guard event collection).

For a finite covariate-cell support and a nonnegative integer , define to be the full finite collection of guard events:

⊢ Lean
Definition 26 [synth_2] (Empirical event frequency).

For observed-data records , each containing a covariate cell, binary instrument, treatment, selection indicator, and optional observed outcome, the empirical frequency of any observable event at sample size and state is

⊢ Lean
Definition 27 [synth_3] (Multiplier process).

For observations , multiplier inputs , sample point , and an observed-data cell , with , , and when selected and when unobserved, the multiplier process is The finite guard-event index collection consists exactly of the cell events indexed by , the arm events indexed by , and the selected-outcome events indexed by .

⊢ Lean

The event collection is finite because both the covariate and outcome supports are finite. Empirical frequencies control every sample input to the capacity construction, while the multiplier process records the local face diagnostics associated with simultaneous max, min, and projection operations.

The construction below names three objects that play different roles, and it is worth separating them before reading it. The first is the reported interval: it is implementable from , , , , , and the support dimensions , which between them determine and . Everything an analyst computes is in that list. The second is , the maximal deviation of the empirical frequencies from over . It depends on the unknown , so it belongs to the analysis rather than to the computation: it is the quantity the concentration argument bounds, while the padding radius actually applied is computed from alone. The third is the multiplier face envelope, a diagnostic for the local nonsmooth behaviour of the endpoint map: it reports which threshold cuts and projection faces are near-active, and it is read on its own, alongside the reported interval. The coverage guarantee rests on the deterministic padding; uniform multiplier calibration under changing survivor support remains open, and is posed in Remark 1.

Algorithm 3 [def:face-aware-inference-handle] (Guarded confidence set ).

Given , , confidence level , instrument-overlap constant , aggregate survivor-complier lower bound , and the finite observable event collection , define the guarded confidence set by the following steps.

  1. Use the retained-cell rule

  2. Define the near-active-face localization tolerance by The near-active lower and upper threshold cuts, the active face of , the active face of , and the coordinatewise nonnegativity faces generated by determine a computable finite diagnostic face envelope for the multiplier process at tolerance .

  3. Let be the maximal empirical deviation over :

  4. Define the union-bound empirical-deviation threshold by

  5. Define the capacity-and-screening deterministic error envelope and deterministic endpoint-padding guard by

  6. Output the conservative guarded confidence set

⊢ Lean

The confidence set in Algorithm 3 uses the deterministic deviation threshold to form a capacity-and-screening envelope . Dividing by the aggregate-mass lower bound yields the endpoint guard , and padding the plug-in endpoints gives the interval . The multiplier component supplies a face-aware diagnostic for nonsmooth local behavior; the stated confidence set is the deterministic guarded interval.

The first large-sample result fixes one observed-data law and characterizes the local endpoint limit on the recovered positive-survivor support.

Theorem 8 [thm:branch-free-pointwise-directional-limit] (Branch-free endpoint limit).

Fix a slate potential-outcome system with finite covariate support and ordered outcome support of size . Let be its observed-data law, let denote the observed cells, let be the weak selection-monotonicity direction, let be the instrument-overlap constant, and let be the screening threshold from Algorithm 2. Suppose that:

  • (Sampling.) For every , the first observations satisfy Assumption 10 with common law .

  • (Overlap.) The instrument propensity satisfies Assumption 5 with constant .

  • (Structural restrictions.) The latent law satisfies Assumptions 1, 2, 3, 4, 6, 7, and 8, with weak selection monotonicity in direction .

  • (Screening rate.) For every , , with and .

  • (Limit inputs.) The finite multinomial central limit theorem applies to the empirical observed-cell vector, and the Hadamard directional delta method applies to tangentially directionally differentiable maps.

Let be the observable capacity vector from Definition 14. Let where is the exact survivor-complier mass in cell , and let . Let be the endpoint map identified by Theorem 1, and let be the projected screened plug-in estimator from Algorithm 2. Then and Moreover, there exist a Gaussian law on observed-cell score functions and a map such that and, for observed cells , The map is the Hadamard directional derivative, at the observed-law mass vector and on the recovered support , of the reduced-support endpoint map. With this derivative,

⊢ Lean

Theorem 8 establishes consistency of the screened support and of the endpoint estimator, then identifies the fixed-law directional limit. The Gaussian input is the centered multinomial limit for the observed cells, and the derivative is taken after the positive-survivor support has been recovered. The branch-free threshold representation handles simultaneous max and min faces through a Hadamard directional derivative; this derivative is the econometric object governing nonsmooth endpoint asymptotics (Fang et al., 2019; Hong et al., 2018).

The next three algebraic facts record how the cellwise masses and threshold cuts scale with nonnegative cell weights. They justify treating covariate-cell probabilities as outer weights in the aggregate endpoint formula and in the empirical analogue.

Lemma 1 [lem:weighted-capacity-mass-homogeneity] (Weighted mass homogeneity).

Let be a capacity array as in Definition 14, and let be real cell weights satisfying for every covariate cell . Define the weighted capacity array cellwise by If denotes the exact survivor-complier mass computed from by the formulas in Definition 14, then, for every covariate cell ,

⊢ Lean

The proof is deferred to Section B.

Lemma 2 [lem:weighted-capacity-lower-cut-homogeneity] (Weighted lower cut homogeneity).

Let be a capacity array, and let be a cell weight function satisfying for every cell . Let be the cellwise weighted capacity array with lower and upper coordinates If denotes the lower threshold-cut value of Definition 16 computed from , and denotes the same lower threshold-cut value computed from , then, for every cell ,

⊢ Lean

The proof is deferred to Section B.

Lemma 3 [lem:weighted-capacity-upper-cut-homogeneity] (Upper cut homogeneity).

Let be a capacity array, and let satisfy for every . Form the cellwise weighted capacity array by Writing for the upper threshold cut of Definition 16 computed from , and for the same cut computed from , one has, for every ,

⊢ Lean

The proof is deferred to Section B.

Lemmas 1, 2, and 3 show that multiplying all capacities in a cell by the same nonnegative weight multiplies the exact mass and both benefit cut values by that weight. Consequently, aggregation over cells preserves the threshold-cut form of the endpoint numerators and denominator.

For the uniform statement, laws are indexed by . The empirical law and auxiliary multipliers are defined law by law, while the deterministic guard keeps the same finite-event form.

Definition 28 [synth_19] (Empirical observed-data law for the -indexed sample).

For each law and sample , the empirical observed-data law is .

Definition 29 [synth_6] (Auxiliary multiplier process).

For each law , let denote auxiliary multiplier variables, conditionally independent given the observed sample and satisfying conditional mean zero and conditional variance one. They generate the diagnostic multiplier process used to form the finite face diagnostic. The guarded confidence interval itself is constructed from the screened endpoint estimator and the deterministic padding .

The empirical law and auxiliary process in Definitions 28 and 29 provide the indexed-sample notation for applying the same construction uniformly over . The multiplier process is diagnostic for the finite face envelope, while containment is delivered by the deterministic padding rule in Algorithm 3.

Theorem 9 [thm:uniform-deterministic-guard] (Uniform deterministic guard).

Fix a family of finite-slate potential-outcome systems indexed by , with nonempty finite covariate support , ordered outcome support of size , and . For each , let be the observed-data law of , where on selected units and is missing off selection. Let be observations taking values in the observed-data cells, let be the auxiliary process used by the guarded interval in Algorithm 3, let be the weak selection-monotonicity direction, and let be the sampling law.

Assume:

  • (Nominal level.) The nominal error level satisfies .

  • (Overlap and mass constants.) The instrument-overlap constant and aggregate survivor-complier mass lower bound satisfy

  • (Screening sequence.) The screening thresholds satisfy for every ,

  • (Uniform law class.) For every law and every , the system belongs to the uniform finite-slate law class of Definition 23 at index , with overlap constant , aggregate survivor-complier mass bound , direction , sampling law , and observations .

Let be the sharp identified interval for the benefit probability from Definition 17, and let be its endpoint map. Let be the projected screened plug-in endpoint estimator from Algorithm 2. Let and be the guarded confidence interval and deterministic endpoint-padding guard from Algorithm 3.

Then, for every , Moreover, where , and

Finally, for every deterministic sequence such that the screening choice satisfies, for every , and the corresponding deterministic guard radius obeys

⊢ Lean

Theorem 9 gives the advertised guarded inference guarantee. For every sample size, the interval from Algorithm 3 contains the full sharp identified interval with probability at least , uniformly over the indexed law family. The same result establishes uniform endpoint consistency under the deterministic guard, with radius ; choosing for any and yields a guard of order . The construction is conservative in the sense of finite-sample containment, and its rate statement shows that the extra screening term can vanish arbitrarily close to the root- scale.

Limitations and future work

The most important practical limitation is the width of the deterministic guard. The guarantee in Theorem 9 is a union bound over the finite event collection , and it buys uniformity at the price of large constants. Taking the smallest configuration the model allows — , , , , and — gives , so before the screening term is added. Since is capped at , the reported interval equals the whole of for every below roughly , and its half-width first falls below — a reported width of about plus the width of the identified interval itself — at of order . At any sample size an applied researcher will encounter, the guarded interval is therefore uninformative: it is a proof that a finite-sample, uniformly valid procedure exists in this model, not a procedure we recommend reporting on its own. The rate is the honest content of the result; the constants are not sharp, and we have not attempted to optimize them. Reducing them — by replacing the union bound over with a self-normalized or empirical-process concentration argument, and by exploiting that the capacity map is -Lipschitz in far fewer directions than suggests — is the most valuable next step for making the guarantee usable, and a simulation study establishing the achievable widths is a prerequisite for any empirical recommendation. We make no claim about practical performance here.

The sharp bounds and guarded confidence set above also leave one calibrated-inference problem for future work. The issue arises when threshold-cut faces intersect, selection-direction signs are close to changing, and the retained covariate support varies with the screening threshold. A useful strengthening of the inference theory would combine the finite deterministic guard in Theorem 9 with a first-order calibration based on the near-active faces recorded by Algorithm 3. The additional sign separation needed for such a direction-stable analysis is the following standard strong variant of the covariate-specific selection-direction condition.

Assumption 11 [ass:direction-margin] (Direction margin).

For every with and ,

⊢ Lean

Assumption 11 is imposed for a strictly positive constant ; read that way it is a sign-separation condition on the covariate-specific selected-complier capacity gap, requiring the selection direction to stay at least away from zero on cells with positive survivor-complier mass. The displayed inequality alone carries no content when , since then holds automatically; positivity of the margin is what the direction-stable analysis below uses, and it is not implied by the inequality as displayed. This is the analogue, for the capacity representation used here, of the strong direction conditions used to stabilize sign-sensitive inference under sample selection (Semenova, 2025). It gives the multiplier program a fixed local direction in cells that contribute to the survivor-complier target, while retaining the threshold-cut ties that make the endpoint map nonsmooth.

Remark 1 [oeq:uniform-face-multiplier] (Uniform face multiplier calibration).

A natural next question is whether the near-active-face directional-multiplier handle admits a calibration such that, uniformly over direction-separated triangular arrays , while allowing arbitrary simultaneous threshold-cut ties and covariate cells whose survivor mass approaches zero and is omitted at the unnormalized threshold .

⊢ Lean

Remark 1 formulates the corresponding open calibration target. The desired confidence set would cover , the sharp identified interval for the survivor-complier benefit probability, uniformly over direction-separated triangular arrays, using , the guarded interval constructed from the finite face diagnostic. The challenge is to preserve first-order calibration when several threshold cuts bind at once and when cells enter or leave the retained support as their empirical survivor mass crosses . This problem is closely connected to modern inference for partially identified treatment effects and selected samples, where uniformity at boundary and contact points is central (Lee et al., 2024; Bartalotti et al., 2023).

Several empirical and modeling extensions are also natural. The Job Corps application that motivates the ordinal wage target can be revisited with the sharp survivor-complier bounds and the guarded interval reported jointly. Continuous outcomes can be handled through discretizations or through an analogue of the threshold-transport construction on richer ordered supports, with the estimand and support conditions stated directly for that setting. Sensitivity analysis for weak selection monotonicity would index departures from Assumption 7 and trace their effect on the endpoint formulas in Definition 17. Additional concordance restrictions, such as restrictions tying potential wage ranks across treatment states, would sharpen the cellwise transport polytopes in Definition 15 and produce a different identified interval whose sharpness would need to be established under the enriched model.

Appendices

Technical derivations and verification note

The appendix collects the supporting derivations for the sharp survivor-complier benefit bounds, the computational construction, and the guarded inference result. The first part works through the capacity algebra behind Definition 16 and the residual completion used by Algorithm 1. The second part records the comparison argument behind Theorem 4. The final technical part gives the concentration calculations that enter the deterministic guard in Algorithm 3 and Theorem 9.

Threshold-cut algebra and residual completion

The threshold-cut calculations organize the cellwise problem in Definition 15 by ordered lower and upper outcome levels. For the lower endpoint, the relevant comparisons aggregate lower-arm mass weakly below a threshold and upper-arm mass weakly below the same threshold; for the upper endpoint, they compare lower-arm mass below a threshold with upper-arm mass above it. These finite ordered-support identities are the algebraic content behind Definition 16 and explain why the endpoint formulas in Definition 17 depend only on observed capacities and cell probabilities.

The residual completion step fills the remaining rows and columns after the threshold mass has been assigned. Its role is to turn the cut values into feasible cellwise couplings in the exact-mass transport set, preserving the row and column inequalities while attaining the desired strict-benefit mass. This is the constructive bridge from the formulas in Definition 16 to the endpoint-attaining laws in Theorem 3. The same finite-dimensional structure also underlies the sparse implementation summarized by Theorem 7.

No-selection reduction

The no-selection reduction specializes the general selected-survivor construction to the setting in which the selection constraint is inactive. Under the hypotheses of Theorem 4, the cellwise capacity arrays collapse to the ordinary complier outcome margins, and the same threshold formulas recover the established ordinal benefit bounds. This comparison fixes the relationship between the survivor-complier estimand studied here, sharp complete-marginal ordinal-outcome benefit bounds (Lu et al., 2018), and selection-bounds analyses of intensive-margin effects under endogenous treatment (Dong et al., 2026).

Guard calculations

The inference appendix derives the deterministic inequalities used by the guarded confidence set in Algorithm 3. The calculations combine finite-support empirical deviations over the event collection in that construction with the screening rule in Algorithm 2. Standard empirical-process and delta-method arguments for finite-dimensional laws supply the pointwise Gaussian approximation underlying Theorem 8 (van der Vaart, 1998). The guard in Theorem 9 then converts uniform control of the observed event probabilities into endpoint padding for the full identified interval.

The face-aware diagnostic records the threshold cuts that are near active at the empirical law. This finite diagnostic is useful because the endpoint map is piecewise linear in the capacity vector, with different linear pieces selected by the active threshold comparisons. The deterministic guard separates two tasks: it bounds capacity and screening errors by an explicit envelope, and it expands the resulting endpoint interval by a finite-sample padding term. This organization follows the econometric logic of inference for nonsmooth and partially identified parameters, where contact sets and binding inequalities determine the local approximation (Fang et al., 2019).

Verification note

The formal assumptions, definitions, propositions, algorithms, theorem statements, and proof dependencies in the paper are machine-checked in Lean 4 for the finite ordered-outcome model stated in the main text. The checking covers the capacity-identification formulas, the threshold-cut endpoint characterization, endpoint attainment by the threshold-flow construction, the no-selection reduction, the three-level witness, the sparse threshold-flow calculation, the pointwise directional limit statement, and the deterministic guarded-inference guarantee.

Scope, toolchain, and commit.

The development is checked with Lean 4, toolchain leanprover/lean4:v4.33.0, at commit 8f179a26a191, in the module tree CausalSmith/PartialID/PID_SlateBenefitPartialtransport_Research. This paper displays numbered formal objects. Of these, are matched to the Lean development in that tree, are presentation-level restatements carrying no Lean declaration of their own (listed below), and — the direction margin of Assumption 11 — is stated only to pose the open calibration problem and is used by no verified theorem. Every declaration certifying a matched object is free of sorry and of added axioms beyond the ambient Lean and Mathlib development.

Objects with no matching declaration.

The seven displayed definitions above are presentation-level restatements introduced for readability rather than separate Lean declarations: the slate-benefit partial-transport system (Definition 11), the threshold-cut capacity aggregates (Definition 18), the cellwise benefit-mass functional (Definition 19), the costed threshold-flow routines (Definition 21), the sparse threshold-flow operation counts (Definition 22), the empirical observed-data law for the indexed sample (Definition 28), and the auxiliary multiplier process (Definition 29). Each names or aggregates quantities that the matched declarations define; none carries a verified claim of its own, and no theorem in this paper is certified through one of them. Readers should treat the verification claim as attaching to the matched statements and their proofs, not to these expository restatements.

Cited inputs.

Published econometric and probability results invoked through citations are bibliographic dependencies; their source proofs are used in the ordinary way through the cited literature. Where a displayed result depends on such an input, the result carries a footnote naming the cited conclusion and stating that the portions depending on it are certified conditional on that conclusion.

Proofs of the main results

Proof of Theorem 1.

Throughout, conditional probabilities on a zero-mass conditioning event are interpreted by the totalized convention as zero; on positive-mass conditioning events they are the usual ratios. Write Cell probabilities are nonnegative because . Now fix with . For the conditional instrument propensity , the totalized convention is the usual ratio on this cell, so By Assumption 5, Hence . Since the two instrument fibers partition the cell, and therefore Thus for both .

For supported , the observed lower-arm contrast pulls back along the observed-data law to Similarly, The observed-data law is the pushforward of the latent law by the observed datum map, whose coordinate events are measurable; hence these are exactly the corresponding conditional probabilities under the latent probability space.

Using Assumptions 1, 2, 3, and 4, for each , and Substitution gives and

By Assumption 6, , so up to null sets. Therefore

The same monotonicity gives and hence

These identities also show nonnegativity in supported cells. If , then for both , so every conditional term in the observable contrasts is zero and the corresponding capacities are nonnegative. Thus the observed-data law is a probability law and its observable capacity contrasts are valid, so it is compatible with the maintained capacity domain.

Because is finite, summing the lower selected-complier identities over outcome levels gives

Applying the same finite partition to the upper selected-complier identities gives

By Definition 14,

It remains to identify . Define the finite restricted measure by for every measurable event . If , all selected-complier masses in the cell are zero. Otherwise, Assumption 7 applies on this restricted cell. If , then -almost surely, so and If , then -almost surely, so and Combining these alternatives with from Definition 14 yields

The same two alternatives determine every strict sign of the gap. If , then , so . Hence forces . If , then , so . Hence forces .

Finally suppose and . The gap identity gives Equivalently, . If , both direction inequalities are vacuous in the complier cell. If , Assumption 7 gives one almost-sure inclusion between and under ; equality of their finite measures upgrades that inclusion to Thus, in either case, both inequalities and hold -almost surely in the cell. Define the updated direction maps by Away from , the original weak-monotonicity condition is unchanged; at , the preceding alternatives supply either direction. Therefore the same latent law, with direction map , witnesses observational admissibility of , and the same latent law, with direction map , witnesses observational admissibility of . The observed-data law is unchanged in both witnesses.

Proof of Theorem 2.

Fix a cell with , and let be the observable capacity vector from . By Theorem 1, is valid, so for all , and Also, by Definition 14, For a coupling array , write These totals satisfy

  1. We first prove the face equivalence. If , then Definition 15 gives When , Each summand is nonnegative, hence Similarly, when ,

    By the cellwise output of Algorithm 1, choose a feasible exact-mass coupling Assume If , the preceding row-exactness argument puts in . The displayed face equalities then put in . Therefore using exact rows for the first equality and exact columns for the second. If , the same argument with columns first puts in , then in , and again gives Thus .

    Conversely, suppose . Then If , then its exact row constraints give so . Since , the column-exactness argument above gives and hence . The reverse inclusion follows symmetrically from exact columns and row-exactness. Therefore The same reasoning sends every element of to , while is immediate from the defining constraints. Hence This proves

  2. Now assume . Since , conditional probabilities given use the denominator . The displayed identity from Theorem 1 yields If , then both are zero. If , then Assumption 7 gives one of the two almost-sure orders on the complier cell: The two indicators have equal conditional integrals by the preceding display, so the ordered indicators are equal almost surely on . Therefore

  3. Under the same zero-gap condition, the argument in the first step also gives Indeed, every has exact rows and exact columns because , and every has bounded rows, bounded columns, and total mass . Combining this equality with the already proved and with from Definition 15, gives

  4. Finally fix a threshold . The prefix sums and , and the upper tail , are those of Definition 16. The lower tail appearing in the statement is Because is finite and each index lies in exactly one of the two sets and , The same finite partition gives When , , and hence

The four displayed conclusions are exactly the asserted cellwise conclusions for every with .

Proof of Theorem 3.

Let be the observable capacity vector from Definition 14, and let By Theorem 1, the capacity vector is valid, the weights satisfy , and, for every supported cell , It also gives the direction implications The capacity-identification step applies on supported cells. On cells with , the observed cell event has zero probability. Hence, for each instrument arm , With the paper’s zero-denominator convention for conditional probabilities, the two conditional probabilities entering every lower-arm contrast and every upper-arm contrast are therefore zero. Thus for every . Since it follows that

Choose as baseline the original compatible law generating . The following cell quantities use the finite-table totalization convention: if , they are the ordinary conditional probabilities given ; if , they are the corresponding zero-cell baseline table entries. Write the baseline complier mass as and write its never-taker and always-taker components as For , the selected-complier masses are bounded by the complier mass, hence For , both capacity totals are zero, so the same inequality holds. Thus the baseline is compatible with and , has conditional complier mass at least in every cell, and carries the displayed never-taker and always-taker components.

Since , fix the reference outcome For each cell, apply the lower and upper threshold-flow constructions from Algorithm 1. Let be the two survivor-complier couplings. They satisfy, for , and their benefit masses are

For each , define the one-sided selected-complier residuals by and the never-selected complier mass by Then, for every cell and every , Indeed, when or , the corresponding identity is the definition of the residual. When , the row upper bounds have total , so every row bound binds; when , the column upper bounds have total , so every column bound binds.

Build, for each , a cell table For , set For , set The nonnegative entries sum to one in every cell: in supported cells the survivor, one-sided, and never-selected complier entries sum to , while Assumption 6 sets the defier stratum to zero, so the never-taker, complier, and always-taker components exhaust the principal-stratum partition; in unsupported cells the table is a point mass.

Let be the law obtained by drawing with probabilities , drawing conditionally on with the same instrument propensity as , and drawing conditionally on according to , independently of given . The associated slate sets The conditional product construction gives conditional instrument independence. The displayed definitions of , , and give treatment consistency, selection consistency, and outcome consistency. The entry with has zero mass in every cell of ; hence the normalized cell table has no defier mass, and the induced slate satisfies treatment monotonicity.

The weak selection direction is also preserved. If and , the implication above gives as impossible, hence so the complier stratum with has zero mass. If and , then is excluded and so the complier stratum with has zero mass. Unsupported cells have , so they contribute no conditional positive-mass case.

The observed law is unchanged. For supported cells, the selected-complier margins displayed above give exactly the observable capacities and ; the never-taker and always-taker margins are the baseline ones; and the complier total is For , put Assumption 5 gives whenever . For treatment state , selection state , and reported outcome , define and define the original latent tail margin For the constructed table, define the decoded observed margin For selected reported outcomes, the table definitions and the displayed selected-complier margins give and The corresponding original latent margins have the same four values: Theorem 1 identifies and with the selected complier outcome masses on supported cells, while Assumption 6 makes the defier terms in the mixed treatment strata equal to zero. Hence for every selected reported outcome . The invalid observed-outcome margins also agree: It remains only the unselected reported outcome. The constructed table has the same conditional mass for each principal treatment event as the baseline, so The original latent tails form the same finite partition, Together with the selected-margin identities, this gives Thus, for every supported cell and every observable singleton, The first equality is the decoded-table construction; the last equality uses Assumptions 1, 2, 3, and 4 and the positive arm mass supplied by Assumption 5. Cells with have zero observed mass under both laws. Hence

Because the observed law is the same, the inequalities in Assumption 5 and the mass condition in Assumption 8 transfer unchanged.

It remains to compute the target value. In supported cells, the survivor coupling of is ; in unsupported cells its contribution is multiplied by . Hence The benefit numerators are Since , division by and the endpoint formula in Algorithm 1 give Thus and are endpoint witnesses.

Finally, the realization statement follows cell by cell. If , the conditional table of equals the lower completion components just defined, and the conditional table of equals the upper completion components. If , the zero-cell baseline convention gives and all baseline never-taker and always-taker component masses in that cell are zero. Therefore the lower and upper threshold-flow completions are both the zero completion, for all . The realization identities in a zero-probability cell are evaluated under the same zero conditioning-cell mass convention: every conditional baseline component used by the completion is zero, and the zero-flow rule sets the lower and upper threshold completions to the zero array. Thus the survivor, one-sided selected-complier, never-selected, never-taker, and always-taker component functionals of and are the zero functionals displayed above. Hence realizes the lower threshold latent completion and realizes the upper threshold latent completion for every covariate cell.

Proof of Theorem 4.
  1. Fix a covariate cell with , and form the observable capacities from . By Theorem 1, these capacities are valid and, in this supported cell, The no-selection submodel implies the two selected-complier masses equal the complier mass: Indeed, if , then and the event is contained in each of and , while both are contained in . Thus both conditional probabilities are squeezed to . If , the same containment in gives both selected-complier masses equal to zero. Substituting these two identities in the displayed formula for gives

  2. It remains to identify the cellwise feasible set. Since , the zero-gap identity gives The exact survivor-complier mass is For any , Definition 15 gives Summing the row inequalities gives so every row inequality binds. The same argument with columns gives so every column inequality binds. Hence . Conversely, every has both exact margins, and therefore has total mass and satisfies the defining inequalities of . Thus

  3. Now assume . By Definition 16, Using , this becomes Division by the positive number commutes with the finite maximum: for any finite nonempty set of real numbers , Applying this to the displayed maximum, with the zero candidate included as , gives

  4. Similarly, Definition 16 gives Division by commutes with the finite minimum: The cap contributes , and the remaining candidates scale termwise. Therefore These two normalized expressions are the fixed-marginal ordinal strict-benefit formulas recorded by (Lu et al., 2018).

Proof of Theorem 5.

Let be the observable capacity vector associated with . By Theorem 1, is a valid capacity vector, for every , and the positive aggregate survivor condition gives For a coupling , write

  1. Fix a cell with . If is any feasible full law with observed law , its survivor-complier coupling is nonnegative. Its row sums are bounded by the lower-arm selected-complier capacities, its column sums are bounded by the upper-arm selected-complier capacities, and its total mass is the exact survivor-complier mass: Thus every feasible full law projects into .

    Conversely, take any . Define a cellwise family of couplings by where is the lower threshold-flow coupling from Algorithm 1. By Theorem 7, for every , and the chosen lies in by hypothesis. Hence We record the realization step as a proof-local realization claim. Let be any family satisfying Choose one compatible generating law for . For each supported cell , put and choose a fixed outcome level . Define the one-sided residual masses and These quantities are nonnegative on every supported cell. The nonnegativity of and follows from the row and column inequalities defining . For , Theorem 1 and Definition 14 give and similarly The equalities summing over outcomes use the finite ordered outcome support, and the inequalities use inclusion of each selected-complier event in the complier event. Hence . The residual formulas are used on supported cells; cells with are assigned a zero-weight normalized table below.

    The residuals complete to the selected-complier margins. If , then , and the row inequalities in have total slack zero; if , the definition of fills the row slack. Thus, in all cases, Similarly, if , then , and the column inequalities have total slack zero; if , the definition of fills the column slack. Hence The complier masses also add up correctly: Indeed, the left side equals when , equals when , and equals when .

    Construct a latent table in cell by assigning to compliers with and , assigning to compliers with , , and , assigning to compliers with , , and , and assigning to compliers with and . Keep the never-taker and always-taker conditional latent components from . On cells with , choose any normalized latent table, since those cells have zero joint weight.

    Combining these conditional tables with the original cell weights and instrument propensities gives a full law . The preceding displays show that the selected-complier outcome margins are and , the complier total is , and the noncomplier components are unchanged; therefore the observed law is . The construction makes the instrument conditionally independent given , and the observed variables are generated from the potential variables by the consistency and exclusion equations. Treatment monotonicity is inherited from the retained principal strata. The one-sided residuals respect the weak-selection direction because Theorem 1 gives the sign implications for . Thus is feasible, and for every supported cell Moreover, since , the same cell decomposition gives

    Applying this proof-local realization claim to , the realized coupling in cell is . Therefore the projection set of feasible full laws is exactly .

  2. Again fix with . If , then . For any , while each row satisfies . Since all row deficits are nonnegative and their sum is zero, Thus , and the reverse inclusion follows because exact rows imply total mass while retaining nonnegativity and the column bounds. Hence

    If , then . The same argument with columns gives for every , and exact columns conversely give total mass . Hence

    If , then , so both the row and column deficit sums vanish. Consequently for every , and the converse is immediate from the same total-mass identity. Therefore

  3. For every , the threshold-cut inequalities are Indeed, the upper endpoint inequality follows from and the threshold cut cover. Fix . For every benefit pair , at least one of or holds; otherwise , a contradiction. Since is entrywise nonnegative, The row and column bounds in then give Together with the bound by , this gives . For the lower bound, nonnegativity gives . Fix . Entry by entry, where denotes an indicator. If , the two sides are equal because ; if and , the left side is zero; if and , the left side is nonpositive; and if and , the left side is zero. Nonnegativity of handles the nonpositive and zero cases. Summing this entrywise comparison gives When , the row sums bind by the preceding step and , so When , the column sums bind and the total row slack is The prefix row slack is at most this total slack, and therefore Combining the two cases gives for every , and hence .

    For the lower and upper threshold-flow couplings from Algorithm 1, Theorem 7 gives Thus

  4. If , then any has and all entries are nonnegative. Hence every entry is zero, so . Conversely, the zero coupling is nonnegative, has all row and column sums equal to zero, and has total mass zero. Since the capacity vector is valid, and for every , so the row and column bounds hold; the total-mass constraint holds because . Thus , and therefore

  5. Define The cellwise endpoint inequality follows by applying the preceding bounds to , and gives

    Let be any feasible full law with observed law . Equality of observed laws transfers the cell weights: For supported cells, the projected coupling lies in ; for unsupported cells, . Therefore and the denominator of the benefit probability is The aggregate identity for the benefit probability gives so every feasible value belongs to .

    For the reverse inclusion, take a real number . If , choose the family . The proof-local realization lemma gives a feasible law with If , set and define Since nonnegativity, the row bounds, the column bounds, and the exact total-mass identity are preserved by convex combinations, Linearity of gives The proof-local realization lemma therefore yields a feasible law with Hence the feasible values are exactly , which is the interval in Definition 17.

Proof of Theorem 6.

Let be the observable capacity vector obtained from on the singleton covariate space and outcome levels .

  1. Consider the latent allocation from Definition 20, written with the canonical hidden-outcome completion The masses are nonnegative and sum to Let be independent of these potential variables with , and set , , and on selected units. Every latent atom has , , and . The instrument propensity in the only cell is , so the overlap inequalities at are The always-selected complier mass is Thus the law satisfies the structural requirements collected in Definition 13, with the single cell retained and weak selection direction .

    Under , the selected outcome masses are , and the unselected mass is . Multiplying by gives the first four observed masses Under , the selected outcome masses are and the unselected mass is . Multiplying by gives the last four observed masses Hence this three-level slate system realizes , has the displayed encoded full law, and satisfies , , and almost surely.

  2. By Definition 14, for , The second conditional probability is for each , and the denominator is . Therefore

  3. Similarly, for , The second conditional probability is for each , and the denominator is . Hence

  4. Summing the displayed capacities gives Thus Using the threshold-cut formulas in Definition 16, the lower-cut correction is The three prefix differences are The maximum with is therefore For the upper cut, and the strict-prefix/tail sums are Taking the minimum gives

  5. The capacity vector is valid because each displayed coordinate of and is nonnegative. The singleton cell has observed weight and the aggregate survivor-complier mass is Since the observed law is induced by the latent slate system constructed above, it lies in the observed-law domain. By Definition 17,

  6. The hypotheses of Theorem 3 hold for the realized three-level slate system: the structural restrictions and tie-safe survivor model hold by the first step, and by the preceding step. Applying that result to , overlap , and the retained singleton cell gives full laws and , with three-level slate systems, that induce and attain the two endpoint values

Combining the six displayed conclusions gives all assertions of Theorem 6.

Proof of Theorem 7.

Choose Both constants are positive. Fix a finite covariate support , an integer , and nonnegative capacity arrays and . Write . For one cell , let be the fixed sparse cell budget charged by the implementation. Since ,

  1. The actual sparse work in cell is the sum of the threshold-cut scan, the lower nested-graph pass, the upper nested-graph pass, and the two completion passes. The threshold-cut scan uses operations. If and are the current row and column lists, each nested pass performs at most operations, and each completion pass performs at most operations. Initially the row and column lists have lengths and . After either nested pass, the residual row and column lists have total length at most . Hence, for the actual cell count , Consequently This proves the actual sparse bound. The charged sparse count is exactly so the same calculation gives

  2. Since , in particular , and therefore Multiplying the sparse bounds by the nonnegative factor gives The dense actual implementation first performs the sparse computation and then materializes the two dense cell matrices, so its actual count is Using the preceding sparse actual bound, The charged dense count is and the same argument gives

  3. The charged sparse count is which depends only on and . For any other nonnegative capacity arrays on the same and , the charged sparse count is the same sum. The charged dense count adds the same deterministic materialization term: Thus the charged sparse counts are equal across such capacity arrays, and the charged dense counts are equal across such capacity arrays.

  4. Fix a cell . Put for the -point ordered support indexing the row and column capacities in the statement; all sums, extrema, and inequalities among indices below use this finite ordered set. With these arrays, Definition 14 gives Denote the deterministic lower and upper scan outputs by and , defined by and The costed sparse routine stores as its value where are the lower and upper sparse allocation traces. By Definition 16, these scan outputs are the threshold-cut values: Moreover

  5. Let and be the matrices obtained by materializing the sparse traces and : for , For any cellwise coupling , the benefit-mass functional used below is the strict-benefit sum For either the lower nonbenefit primary pass or the upper benefit primary pass, the input row and column lists have lengths and . If is the primary allocation list and , are the two residual lists, the recursive pass gives The complete bipartite residual pass consumes one residual row or one residual column at each emitted allocation and stops when one residual side is exhausted, so Thus and, since is an integer, Materializing a sparse trace can only put positive mass on a listed row-column pair, hence

    Every emitted mass in the two primary passes and in the residual completion pass is the minimum of two nonnegative residual capacities, and each residual update subtracts that minimum. Therefore for all . Let and be the materialized primary and completion parts, so . For the row and column residual lists left by the primary pass, define the residual capacities The primary pass conserves each original capacity after its residual is included, while the completion pass uses at most the residual capacities: and Consequently The completion transports the smaller of the two residual totals, and the primary pass conserves the original row and column totals. Hence By the definition of in Definition 15, these displays give

    It remains to identify the two benefit masses. For the lower pass, let be the total mass emitted by the primary nonbenefit pass. Every primary lower allocation satisfies , and every residual row-column pair left for completion satisfies . Hence the primary part contributes zero benefit mass and the completion part contributes all of its transported mass to benefit edges. Since the completed lower trace transports total mass , For the upper pass, let be the total mass emitted by the primary benefit pass. Every primary upper allocation satisfies , and every residual row-column pair left for completion satisfies . Thus the residual completion contributes zero benefit mass and Define the finite lower and upper cut families on by where . For the lower primary pass, the empty cut bound is For an ordinary threshold , each nonbenefit primary allocation is counted by the row side or by the column side , so Using the primary conservation identities and the nonnegative residual capacities, and therefore Thus The reverse inequality is obtained from the boundary invariant of the same ascending pass. The input row and column lists are ordered increasingly. By induction on the sum of their lengths, the recursive pass maintains and, for every residual row and emitted allocation , The empty-list cases are immediate. In an eligible step , the pass emits and recurses after reducing or deleting the exhausted endpoint; the reduced lists preserve the increasing order, and any later residual row has index at least the emitted row index, so the new emitted allocation cannot cross a later residual-row boundary. In an ineligible step , the row is moved to the residual list; every column that can remain residual has index at least , and every later emitted allocation uses such a column, giving both the residual separation and the displayed noncrossing condition for this new residual row.

    If the lower primary pass leaves no row residual, row-total conservation gives , and the column bound above gives . Thus, using Definition 14, so the infimum is at most . If a row residual remains, let be the largest row index among the lower residual rows. The finite maximum transfers the preceding invariants to the cut: and no primary lower allocation with positive mass has both and . Hence Since every primary lower allocation satisfies and no such allocation crosses the cut , its mass is counted exactly once by the row side or the column side : Adding the zero residual sums and applying row and column conservation gives so again . Consequently For the upper primary pass the cut bounds use the strict-benefit cuts. The empty cut gives and, for every , each benefit primary allocation is counted by the row side or by the column side : Thus . For the reverse inequality, use the symmetric boundary invariant for the descending benefit pass. Its input row and column lists are ordered decreasingly, and induction on the same recursion gives and, for every residual row and emitted allocation , The eligible branch emits and recurses on decreasing lists after the exhausted endpoint is reduced or deleted; any later residual row has index at most the emitted row index, so the new emitted allocation cannot cross its residual-row boundary. In the ineligible branch , the row is moved to the residual list; all remaining columns have index at most , and all later emitted allocations use such columns, giving the residual separation and noncrossing condition for the new residual row.

    If the upper primary pass leaves no row residual, then and the column bound above gives , so, by Definition 14, If a row residual remains, let be the smallest row index among the upper residual rows. The finite minimum transfers the preceding invariants to the cut: and no primary upper allocation with positive mass has both and . Hence Since every primary upper allocation satisfies and no such allocation crosses the cut , its mass is counted exactly once by the row side or the column side : Conservation then yields Therefore

    For the lower endpoint, the partition identity , together with from Definition 14, implies Therefore, by Definition 16, For the upper endpoint, the displayed formula for is exactly the upper threshold-cut family in Definition 16, and hence Since was arbitrary, the asserted cellwise output properties hold in every covariate cell.

Proof of Theorem 8.
  1. Let be the finite set of observed cells and write For an event , set The finite multinomial central limit input applied under Assumption 10 gives a Gaussian law on such that and, for observed cells , Here the displayed empirical vector is the atom vector of the sample empirical law, because its -coordinate is exactly the centered empirical frequency of .

  2. Put For an arbitrary vector , define and form the raw contrasts Let and . From these projected capacities define by the formulas in Definitions 14 and 16, and define the fixed-support endpoint map with the quotient coordinates totalized by the value when the common denominator is zero.

  3. Theorem 1 gives compatibility of the observable capacities, , and the structural identification of . Since Assumption 8 gives the recovered support is nonempty after restriction to positive products. If , then , hence . Writing , Assumption 5 gives , , and . Therefore and Thus all conditional denominators used by on are positive at . Moreover, because at is . The nonnegativity supplied by Theorem 1 then yields and Capacity validity gives and , so Consequently For the last equality, if , then ; since Theorem 1 gives and the preceding display gives , we have , hence .

  4. The map is Hadamard directionally differentiable at along all directions in . Indeed, on the neighborhood where the arm denominators in the preceding step are positive, the raw conditional contrasts are smooth quotient maps. The coordinatewise positive-part map has scalar directional derivative at every nonnegative coordinate . Finite sums preserve directional differentiability, the scalar minimum gives the derivative of , and the finite maximum and minimum defining and give the active-face directional derivatives of the threshold cuts. Therefore the denominator and the two numerators in are continuously Hadamard directionally differentiable at . Since the denominator is positive at , the quotient rule gives a continuous directional derivative satisfying whenever and in . Applying the directional delta-method input to gives

  5. The population value of the fixed-support map equals the identified endpoint map. For every array satisfying one has because implies and hence . The capacity validity supplied by Theorem 1 gives for every . Hence so The lower formula in Definition 16 includes the candidate , and therefore To prove the upper bound for , it is enough to bound every lower-threshold candidate by . The zero candidate is bounded by by the preceding display. For a threshold , put so that and . If , then and because and . If , then , , and Thus . For the upper threshold cut, Definition 16 writes as the minimum of and the threshold sums . The candidate gives , while all candidates are nonnegative by capacity validity, so . Consequently Using the restriction identity with gives where the final equality is the endpoint formula in Definition 17.

  6. For each cell , define the empirical screening score The screening rule in Algorithm 2 is exactly If , the same positive-denominator argument as above and the directional delta method applied to the scalar score map give tightness of the scaled score error: for every , there is such that, for all , If , then , and independent sampling gives

  7. Fix a cell and . If , then the preceding display gives exact agreement of the screened and population membership decisions. Suppose . If , write For all sufficiently large , On the event that is not screened, , and hence If , then . For all sufficiently large , On the event that is screened, , and therefore Thus, for every , eventually. A finite union over yields

  8. On the event the screened sums in are exactly the fixed-support sums defining . The support is nonempty and each retained score exceeds , so the screened denominator is positive and the fallback branch in Algorithm 2 is inactive on this event. Hence Together with , equality with probability tending to one transfers the fixed-support weak limit to the screened estimator:

  9. It remains to record consistency. The weak convergence just proved implies tightness of the scaled endpoint error: for every there is such that, for all , Given , choose large enough that Then and therefore eventually. Since is arbitrary, The displayed Gaussian law, covariance identities, support recovery, consistency, and directional weak limit are precisely the asserted conclusions.

Proof of Lemma 1.

Fix a covariate cell . By Definition 14, the capacity totals formed from the weighted array are

  1. Substituting the definition of into the first weighted total gives where the middle equality is finite-sum linearity.

  2. The same computation for the upper-arm total gives This finite-sum identity also covers the degenerate case in which the outcome index set is empty, since both sums are then zero.

  3. Since the weight satisfies , multiplication by commutes with the binary minimum: Combining this identity with the two displayed total identities yields again using the definition of in Definition 14.

Proof of Lemma 2.

Fix a cell . Write , so . By Definition 16, the lower cut can be written as the finite maximum over the dummy alternative and the threshold alternatives: The same formula applied to gives . Since multiplication by commutes with a finite maximum, Thus it remains to compare the dummy term and each threshold term.

For the dummy alternative the equality is immediate, because the term is on both sides: Now fix . The weighted prefix sums satisfy and similarly The weighted capacity-total gap is Because , Combining these identities, The dummy term and all threshold terms therefore agree with the corresponding -scaled unweighted terms inside the same finite maximum. Hence

Proof of Lemma 3.

Fix . Let , where is an auxiliary marker. For , define the unweighted and weighted upper-cut candidates by and where Thus is nonempty even when there are no threshold indices.

  1. By Definition 16, the upper cut is the minimum over these candidates: For any finite nonempty index set, multiplication by a nonnegative scalar commutes with the minimum: Indeed, if both sides are zero, while if the order-preserving map carries the least element of the finite family to the least element of the scaled family. Hence

  2. It remains to identify the weighted candidates with the scaled unweighted candidates. For the marker candidate, Lemma 1 gives For any threshold , and similarly Therefore Thus for every .

  3. Combining the previous displays, This is the desired identity for the fixed cell , and was arbitrary.

Proof of Theorem 9.
  1. Fix , , and . Let . The finite guard-event collection used here is where Define The estimator in Algorithm 2 uses the totalized conditional ratio Set By Definition 23, the law at satisfies the structural restrictions, the i.i.d. sampling condition, and Assumption 9. Thus Theorem 1 gives valid observable capacities and . The overlap condition gives the arm-mass inequality Indeed, when the inequality is immediate from nonnegativity. When , write . For , Assumption 5 gives For , the binary partition and the upper-overlap bound in Assumption 5 give The uniform aggregate bound gives .

    Suppose and . Fix , , and . Write and The definition of and the bound give Moreover because each selected-outcome event is contained in its arm event. Therefore When , the three displayed deviation inequalities force , , and , so the two totalized conditional expressions are identical, including the case of zero arm mass. Assume now that . If , both totalized conditionals lie in , , and hence where the last inequality uses . If , then the arm-overlap inequality gives Thus the ratio branch is active for both denominators, and The numerator bound follows from and the denominator lower bound is . Since , again using .

    Each raw lower or upper capacity coordinate is a difference of two such conditionals, and the projection in Algorithm 2 is coordinatewise positive part, a one-Lipschitz map against the nonnegative population capacities. Hence Put For each cell define the unconditional population and empirical projected capacity coordinates by Let , , and denote, respectively, the survivor-complier mass and the lower and upper threshold cuts computed from the unconditional capacity array . Let , , and denote the corresponding quantities computed from . Since and , Lemmas 1, 2, and 3 give the homogeneity identities and The preceding coordinate bound says that every lower and upper coordinate in these unconditional arrays changes by at most . Hence each total or changes by at most , and the difference of these totals changes by at most . Since , finite maxima, and finite minima are one-Lipschitz in the sup norm, the lower threshold candidate changes by at most The zero candidate in the outer maximum is unchanged, so changes by at most . For the upper cut, the mass candidate changes by at most , and each strict-prefix-plus-tail candidate changes by at most ; therefore changes by at most . The same argument gives the bound for the unconditional survivor mass . Summing over the finite covariate support yields and Before applying the screening rule, record the cut bounds used for both the population capacities and the projected empirical capacities. Fix a covariate cell and either one of these two nonnegative capacity arrays; write its associated totals, gap, mass, and cuts with a superscript : All coordinates are nonnegative, so . By the formula in Definition 16, is the maximum of the zero candidate and the threshold candidates, and therefore The upper cut is the minimum of and nonnegative strict-prefix-plus-tail sums, so It remains to bound the lower cut above by the mass. For a threshold , define its lower-cut candidate by If , then , , and ; hence If , then and , so because and . The zero candidate is also at most . Thus every candidate in the finite maximum defining the lower cut is at most , and Applying these inequalities first to the valid population capacities and then to the projected empirical capacities gives, for every , Screening removes, in each cell, at most from the empirical mass and at most from each empirical numerator. Indeed, a screened-out cell satisfies , while the just-proved empirical cut bounds give and . Thus, with Since , , and the retained-cell indicators are nonnegative, the same cut bounds also imply

    If , then so the ratio branch of is active. For either endpoint numerator and its empirical screened counterpart , If , then , and both the target endpoints and the plug-in endpoints lie in . Combining the two cases gives and the sharper bound

  2. For , the maximal guard deviation has the finite second-moment tail bound To see this, for put The i.i.d. sampling condition gives independent mean-zero variables with Hence Chebyshev’s inequality and a finite union bound over give the displayed tail bound.

    Since is nonempty, contains a cell event, so and therefore

  3. On the event , the first step with and Algorithm 3 give Let . By Definition 17, and the endpoint bounds established from give and . Thus which is exactly . The tail bound in the previous step therefore implies Taking the infimum over proves the finite-sample coverage claim for this , and was arbitrary.

  4. Define Since , , and , we have . Together with the screening-threshold assumption , this gives . Moreover, so eventually For such , if , the sharper deterministic bound from the first step gives Because and , Together with the bound by from the first deterministic inequality, this yields Consequently, eventually in , The tail bound gives, uniformly in ,

    The right-hand side tends to because .

    This proves

  5. Let For , Therefore Algorithm 3 gives which proves

    Now let , , and , and set For every , . Also Since , eventually , and hence . The preceding guard bound then gives, eventually, so

References

  • Angrist, Joshua D. and Imbens, Guido W. and Rubin, Donald B. (1996). Identification of Causal Effects Using Instrumental Variables. Journal of the American Statistical Association. doi
  • Kennedy, Edward H. and Harris, Scott and Keele, Luke J. (2019). Survivor-Complier Effects in the Presence of Selection on Treatment, With Application to a Study of Prompt {ICU} Admission. Journal of the American Statistical Association. doi
  • Chen, Xiaohong and Flores, Carlos A. (2015). Bounds on Treatment Effects in the Presence of Sample Selection and Noncompliance: The Wage Effects of {Job Corps}. Journal of Business \& Economic Statistics. doi
  • Dong, Yingying and Heiler, Phillip (2026). Sharp Bounds and Inference in Sample Selection Models with Treatment Endogeneity. . arXiv
  • Lu, Jiannan and Ding, Peng and Dasgupta, Tirthankar (2018). Treatment Effects on Ordinal Outcomes: Causal Estimands and Sharp Bounds. Journal of Educational and Behavioral Statistics. doi
  • Gabriel, Erin E. and Sachs, Michael C. and Jensen, Annette K. (2024). Sharp Symbolic Nonparametric Bounds for Measures of Benefit in Observational and Imperfect Randomized Studies with Ordinal Outcomes. Biometrika. doi
  • {de Aguas}, Johan and Krumscheid, Sebastian and Pensar, Johan and Biele, Guido (2025). The Probability of Tiered Benefit: Partial Identification with Robust and Stable Inference. Proceedings of the First Conference on Causal Learning and Reasoning. arXiv
  • Fang, Zheng and Santos, Andres (2019). Inference on Directionally Differentiable Functions. Review of Economic Studies. doi
  • Chernozhukov, Victor and Lee, Sokbae and Rosen, Adam M. (2013). Intersection Bounds: Estimation and Inference. Econometrica. doi
  • Hong, Han and Li, Jessie (2018). The Numerical Delta Method. Journal of Econometrics. doi
  • Semenova, Vira (2025). Generalized {Lee} Bounds. Journal of Econometrics. doi
  • Noack, Claudia (2026). Sensitivity of {LATE} Estimates to Violations of the Monotonicity Assumption. . arXiv
  • Fan, Yanqin and Park, Sang Soo (2010). Sharp Bounds on the Distribution of Treatment Effects and Their Statistical Inference. Econometric Theory. doi
  • Russell, Thomas M. (2021). Sharp Bounds on Functionals of the Joint Distribution in the Analysis of Treatment Effects. Journal of Business \& Economic Statistics. doi
  • {van der Vaart}, Aad W. (1998). Asymptotic Statistics. Cambridge University Press. doi
  • Chen, Xiaohang and Li, Fan (2026). Principal Stratification with {U}-Statistics Under Principal Ignorability. Journal of the Royal Statistical Society Series B: Statistical Methodology. doi
  • Imbens, Guido W. and Angrist, Joshua D. (1994). Identification and Estimation of Local Average Treatment Effects. Econometrica. doi
  • Balke, Alexander and Pearl, Judea (1997). Bounds on Treatment Effects from Studies with Imperfect Compliance. Journal of the American Statistical Association. doi
  • Frangakis, Constantine E. and Rubin, Donald B. (2002). Principal Stratification in Causal Inference. Biometrics. doi
  • Heckman, James J. (1979). Sample Selection Bias as a Specification Error. Econometrica. doi
  • Lee, David S. (2009). Training, Wages, and Sample Selection: Estimating Sharp Bounds on Treatment Effects. Review of Economic Studies. doi
  • Manski, Charles F. (1990). Nonparametric Bounds on Treatment Effects. American Economic Review.
  • Manski, Charles F. (1997). Monotone Treatment Response. Econometrica. doi
  • Manski, Charles F. and Pepper, John V. (2000). Monotone Instrumental Variables: With an Application to the Returns to Schooling. Econometrica. doi
  • Manski, Charles F. (2003). Partial Identification of Probability Distributions. Springer. doi
  • Imbens, Guido W. and Manski, Charles F. (2004). Confidence Intervals for Partially Identified Parameters. Econometrica. doi
  • Chernozhukov, Victor and Hong, Han and Tamer, Elie (2007). Estimation and Confidence Regions for Parameter Sets in Econometric Models. Econometrica. doi
  • Romano, Joseph P. and Shaikh, Azeem M. (2010). Inference for the Identified Set in Partially Identified Econometric Models. Econometrica. doi
  • Tamer, Elie (2010). Partial Identification in Econometrics. Annual Review of Economics. doi
  • Molinari, Francesca (2020). Microeconometrics with Partial Identification. Handbook of Econometrics. arXiv
  • Makarov, G. D. (1982). Estimates for the Distribution Function of a Sum of Two Random Variables When the Marginal Distributions are Fixed. Theory of Probability and Its Applications. doi
  • Williamson, Robert C. and Downs, Tom (1990). Probabilistic Arithmetic. {I}. Numerical Methods for Calculating Convolutions and Dependency Bounds. International Journal of Approximate Reasoning. doi
  • Villani, C{\'e}dric (2009). Optimal Transport: Old and New. Springer. doi
  • Lu, Jiannan and Zhang, Yukun and Ding, Peng (2020). Sharp Bounds on the Relative Treatment Effect for Ordinal Outcomes. Biometrics. doi
  • Tian, Jin and Pearl, Judea (2000). Probabilities of Causation: Bounds and Identification. Annals of Mathematics and Artificial Intelligence. doi
  • Sachs, Michael C. and Gabriel, Erin E. and Sj{\"o}lander, Arvid (2023). A General Method for Deriving Tight Symbolic Bounds on Causal Effects. Journal of Computational and Graphical Statistics. doi
  • Duarte, Guilherme and Finkelstein, Noam and Knox, Dean and Mummolo, Jonathan and Shpitser, Ilya (2023). An Automated Approach to Causal Inference in Discrete Settings. Journal of the American Statistical Association. doi
  • Blanco, German and Flores, Carlos A. and Flores-Lagunes, Alfonso (2013). Bounds on Average and Quantile Treatment Effects of {Job Corps} Training on Wages. Journal of Human Resources. doi
  • Chen, Xiaohong and Flores, Carlos A. and Flores-Lagunes, Alfonso (2018). Going Beyond {LATE}: Bounding Average Treatment Effects of {Job Corps} Training. Journal of Human Resources.
  • Bartalotti, Ot{\'a}vio and K{\'e}dagni, D{\'e}sir{\'e} and Possebom, Vitor (2023). Identifying Marginal Treatment Effects in the Presence of Sample Selection. Journal of Econometrics. doi
  • Levis, Alexander W. and Bonvini, Matteo and Kennedy, Edward H. and Keele, Luke J. (2025). Covariate-Assisted Bounds on Causal Effects with Instrumental Variables. Journal of the Royal Statistical Society Series B: Statistical Methodology. doi
  • Ben-Michael, Eli (2025). Partial Identification via Conditional Linear Programs: Estimation and Policy Learning. . arXiv
  • Khandamiryan, Gevorg and Semenova, Vira (2026). Adaptive Estimation of Aggregated Values of Conditional Linear Programs. . arXiv
  • Lee, Ying-Ying and Liu, Chu-An (2024). {Lee} Bounds with a Continuous Treatment in Sample Selection. . arXiv
  • Possebom, Vitor and Riva, Flavio (2025). Probability of Causation With Sample Selection: A Reanalysis of the Impacts of {J{\'o}venes en Acci{\'o}n} on Formality. Journal of Business \& Economic Statistics. doi