Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

PUATE: Efficient Average Treatment Effect Estimation from Treated (Positive) and Unlabeled Units

T0 review · 3 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read This paper claims that average treatment effect (ATE) estimation from treated-positive and unlabeled data has an attainable semiparametric efficiency bound, and constructs cross-fitted estimators that attain it in both censoring and…

desk verdict Careful EIF derivation for PU-style ATE, but the headline efficiency claim only holds with known censoring propensity; the abstract overstates it. read the letter →

arxiv 2501.19345 v2 pith:XAMYEQEJ submitted 2025-01-31 cs.LG econ.EMmath.STstat.MEstat.MLstat.TH

classification cs.LGecon.EMmath.STstat.MEstat.MLstat.TH MSC 62G2062D1062G05
keywords averagetreatmenteffectpositiveandunlabeledlearningsemiparametricefficiencyboundefficientinfluencefunctionmissinglabelscensoringsettingcase-controlcross-fitting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether the average treatment effect can be estimated optimally when the data contain only confirmed treated units and an unlabeled pool whose treatment status is unknown, a positive-and-unlabeled (PU) version of causal inference. It derives the semiparametric efficiency bound for two sampling designs, the censoring setting and the case-control setting: the bound is the smallest asymptotic variance any regular estimator can attain. It then constructs cross-fitted estimators whose asymptotic variance matches that bound, so the missing treatment labels do not by themselves force a loss of asymptotic optimality. The catch is that this efficiency result requires the censoring propensity score, the probability that an unlabeled unit was actually treated given its covariates, to be known or supplied by an auxiliary dataset; without it, the paper establishes consistency but not $\sqrt{n}$-efficiency.

What carries the argument

The load-bearing object is the efficient influence function (EIF), the score whose variance defines the efficiency bound and whose sample average, with estimated nuisance parameters, gives the optimal estimator. In the censoring setting it is $\Psi_{\mathrm{cens}}=S_{\mathrm{cens}}-\tau_0$, where $S_{\mathrm{cens}}$ combines inverse-probability-weighted residuals of the observed treated outcomes and of the unlabeled-pool outcome, plus a level correction involving the conditional means $\mu_{T,0}$ and $\nu_0$; the censoring propensity score $g_0(d\mid X)=P(D=d\mid X,O=0)$ enters as a known weight. The estimator is the sample average of $S_{\mathrm{cens}}$ with cross-fitted nuisance estimates, so the EIF acts as the estimating equation. For the case-control setting, the corresponding objects are the pair of influence functions $S_{cc}^{(T)}$ and $S_{cc}^{(U)}$, where the case-control propensity $e_0$ and the covariate density ratio $r_0=\zeta_0/\zeta_{T,0}$ are the oracle inputs.

What would settle it

Simulate a DGP satisfying Assumptions 3.2, 3.3, 4.5, and 4.6 with the censoring propensity score known, run the cross-fitted estimator at $n=3000$, and compare the empirical variance of $\sqrt{n}(\hat{\tau}_n^{\mathrm{cens\text{-}eff}}-\tau_0)$ with $V_{\mathrm{cens}}$; Theorem 4.7 fails if the discrepancy exceeds Monte Carlo error. A complementary check is to estimate $g_0$ from the same primary sample without an auxiliary dataset; if $\sqrt{n}$-normality at variance $V_{\mathrm{cens}}$ still holds, the paper's claim that the oracle $g_0$ is needed is contradicted.

Watch

Extended reading notes

Core claim

The paper's central claim is that the ATE in the PU setting has an attainable semiparametric efficiency bound. In the censoring setting, the efficient influence function is $\Psi_{\mathrm{cens}}(X,O,Y;\mu_{T,0},\nu_0,\pi_0,g_0)=S_{\mathrm{cens}}-\tau_0$, and the efficiency bound is $V_{\mathrm{cens}}=\mathbb{E}[\Psi_{\mathrm{cens}}^2]$; Theorem 4.7 shows that a cross-fitted estimator formed by averaging $S_{\mathrm{cens}}$ with estimated outcome regressions and observation probabilities, but with the censoring propensity score $g_0$ known, satisfies $\sqrt{n}(\hat{\tau}_n^{\mathrm{cens\text{-}eff}}-\tau_0)\xrightarrow{d}N(0,V_{\mathrm{cens}})$. In the case-control setting, Theorem D.7 gives the analogous result with the case-control propensity score $e_0$ and density ratio $r_0$ known. The paper further shows that when $g_0$ is estimated from the same data without an auxiliary sample, it cannot establish $\sqrt{n}$-consistency; with an auxiliary covariate-only sample, Corollary 4.9 restores asymptotic normality.

Load-bearing premise

The load-bearing premise is that we know the probability that an unlabeled unit actually received the treatment, given its covariates, or can estimate that probability from a separate auxiliary dataset; if that probability is estimated only from the primary data, the paper's efficiency claim does not go through.

Editorial extensions

If this is right

  • If Theorem 4.7 holds, the proposed censoring-setting estimator is asymptotically optimal: its variance equals the efficiency bound, so no regular estimator using the same information can have smaller asymptotic variance.
  • If Theorem D.7 holds, the case-control estimator attains its bound using only consistent outcome regressions once $e_0$ and $r_0$ are known; the nuisance estimation burden is lighter than in the censoring setting.
  • With consistent $g_0$, consistency of the censoring estimator is doubly robust: it survives if either the observation model $\pi_0$ or both outcome regressions $\mu_{T,0}$ and $\nu_0$ are consistent.
  • The IPW estimator has asymptotic variance at least $V_{\mathrm{cens}}$, so the efficient estimator strictly improves on it whenever the outcome means are nonzero.
  • When $g_0$ is estimated from the same primary data without an auxiliary sample, the paper's asymptotic-normality guarantee does not apply; practitioners are left with consistency, not efficiency.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implication the paper leaves implicit is that the main practical obstacle is assembling the auxiliary covariate-only data needed to estimate $g_0$; the paper does not analyze how large $n_{\mathrm{aux}}$ must be relative to $n$ to make the efficiency approximation tight.
  • A natural testable extension is to derive the exact asymptotic variance when $g_0$ is estimated from an auxiliary sample of size $n_{\mathrm{aux}}$ and to show how the efficiency bound is approached as $n_{\mathrm{aux}}\to\infty$.
  • The censoring setting's oracle assumption suggests a partial-identification reading: without a reliable $g_0$, the ATE may still be bounded rather than point-estimated, and connecting this to existing bounds for missing-treatment data would be a direct follow-up.
  • In the case-control setting, the density ratio $r_0$ is put on the same oracle footing as $e_0$; since direct density-ratio estimation is often harder than propensity estimation, whether plug-in $r_0$ estimators preserve efficiency is an open question this paper does not settle.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies estimation of the average treatment effect when the data consist of a treated (positive) group and an unlabeled group whose treatment status is unknown, a setting the authors call PUATE. Two sampling schemes are considered: a censoring setting with a single dataset containing observed treated units and units with missing treatment labels, and a case-control setting with separate treated and unlabeled samples. The authors derive efficient influence functions and semiparametric efficiency bounds for both settings, propose cross-fitted estimators based on those influence functions, and prove consistency and asymptotic normality under rate conditions on nuisance estimators. The central theoretical results are Theorem 4.7 for the censoring setting and Theorem D.7 for the case-control setting. The paper also reports simulations and semi-synthetic IHDP experiments.

Significance. If the results are read as an oracle-efficiency theory, the paper makes a useful contribution: it explicitly derives the efficient influence functions for a PU-type missing-treatment problem, constructs cross-fitted estimators, and provides the corresponding asymptotic normality results. The paper is also transparent about its main limitation, namely that asymptotic normality requires the censoring propensity score (or the case-control propensity score and density ratio) to be known or estimated from an independent auxiliary dataset. I found no circularity in the derivations; the issue is instead one of identification relative to the model stated in the abstract. Once the claims are re-framed as oracle-model efficiency results, the technical machinery is informative and the simulation study, though limited, is a reasonable check on the finite-sample behavior of the proposed estimators. The uncorrected efficiency bound in Theorem 4.2 and the overstatement of the efficiency claim in the abstract are the main obstacles to publication in the current form.

major comments (3)
  1. [Abstract; §4.1 (Lemma 4.1, Theorem 4.2); Appendix H; §4.4-4.5] The efficiency claim is stated for the model described by Assumptions 3.2–3.3, but Appendix H derives the efficient influence function in parametric submodels that keep g0(d|X) fixed: the submodel density is p_etilde(y|x;θ) = g0(1|x)p_Y(1)(y|x;θ) + g0(0|x)p_Y(0)(y|x;θ). Therefore the tangent space contains no scores for the censoring propensity score g0, and V_cens is the efficiency bound for an oracle model in which g0 is known, not for the PUATE model in which g0 is an unknown component of the data-generating process. In the latter model τ0 is not identified: for any observed law of (X,O,Y), one can change g0(1|X) and redefine p_Y(0) so that the mixture q(y|x) = g0(1|x)p_Y(1)(y|x) + (1-g0(1|x))p_Y(0)(y|x) is unchanged while E[Y(0)] changes. Consequently no regular √n estimator exists in that model, and the abstract's sentence 'We then construct semiparametric efficient ATE estimators that attain these bounds' is not supported. Theorem 4.7 explicitly requires Assumption 4.5 (known g0), and Section 4.5 concedes that √n-consistency cannot be obtained when g0 is estimated from the same data. The same issue applies to the case-control setting, where Assumption D.5 fixes e0 and r0. I recommend re-stating Lemma 4.1, Theorem 4.2, and the case-control analogues as oracle-model results, or as results with an independent auxiliary dataset (Corollary 4.9), and revising the abstract accordingly.
  2. [§4.1, Theorem 4.2] The displayed variance formula in Theorem 4.2 is inconsistent with the efficient influence function in Lemma 4.1. Writing S_cens for the influence-function score, the O=1 terms combine with coefficient 1 + g0(1|X)/g0(0|X) = 1/g0(0|X), not (1 - g0(1|X)/g0(0|X)), and the O=0 part has variance Var(eY|X)/π0(0|X), not Var(eY|X)/π0(1|X). The correct expression should be E[ (1/g0(0|X)^2)(Var(Y(1)|X)/π0(1|X) + Var(eY|X)/π0(0|X)) + (τ0(X)-τ0)^2 ]. This error is load-bearing because V_cens is the claimed efficiency bound used in Theorem 4.7, Corollary 4.9, and the simulation discussion. Please correct the theorem and verify all uses of V_cens.
  3. [§5, Theorem 5.1; Appendix D, Theorem D.7] The case-control efficiency claim has the same oracle-conditional structure as the censoring result. Theorem 5.1 and Theorem D.7 assume that e0(d|X) and r0(X) are known (Assumption D.5). Without such knowledge, or without additional identification assumptions such as an identifiable class prior, the unlabeled sample is a mixture with unknown weight e0(1|X), and τ0 is not identified from the stated data-generating process. The abstract presents the case-control result as an unconditional efficiency bound, which is not supported by the formal theorems. Please make the oracle assumption explicit in the main text and abstract for both settings, and state precisely which quantities are assumed known versus estimated.
minor comments (4)
  1. [Appendix I.2, proof of Theorem 4.7] In the displayed calculation of the second-order remainder, the third term in the expression for the conditional expectation appears with bπ_n(0|X) in the denominator where the EIF has bπ_n(1|X), and similarly π0(0) appears where π0(1) is required. This appears to be a typo, but it should be corrected because it is part of the proof of the main asymptotic normality result.
  2. [§2.3] The definition of Yi uses eY_i before eY_i is defined; reorder the two displayed equations for readability.
  3. [§4.1 and §4.2] The phrases 'If Assumptions 3.2–3.3' and 'such a derivation of the efficient estimator as the estimation equation approach' are incomplete; please edit for grammatical completeness.
  4. [Appendix M, Tables 4-5] In the semi-synthetic IHDP censoring experiments, the proposed efficient estimator has coverage ratios of 0.22 (response surface A) and 0.01 (response surface B), far below the nominal level. The paper should comment on this discrepancy and whether the asymptotic approximation is expected to be reliable at n = 747 with estimated g0.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular reasoning identified; the oracle propensity-score assumption is a stated limitation rather than a circular step.

full rationale

The paper's central result—semiparametric efficiency of the proposed ATE estimators in the censoring and case-control PU settings—is derived from standard semiparametric theory (van der Vaart 1998; Hahn 1998) by computing the efficient influence function, not by assuming the estimator's own variance or the target value. The formulas for Ψcens and the bound Vcens are obtained from the score of parametric submodels and the Riesz representation theorem, and the asymptotic normality in Theorem 4.7 is a standard cross-fitting argument. The main caveat, which the paper itself states in Assumption 4.5 and Section 4.5, is that the efficiency and √n-normality results require the censoring propensity score g0 to be known (or estimated from an auxiliary dataset), and the proof of Lemma 4.1 fixes g0 in the parametric submodel. This is an oracle-model limitation and a possible mismatch with the abstract's unconditional wording, but it is not circular: the bound and the estimator are not defined in terms of each other, and the regularity conditions on nuisance estimators are stated independently. Citations to Uehara et al. (2020) and Kato et al. (2024) provide the general stratified-sampling efficiency framework; they do not assume the PUATE efficiency result, so the self-citations are technical and not load-bearing. No step reduces by construction to its own input.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

No free parameters are used in the theoretical results; simulation parameters are generative modeling choices, not part of the method. The central theorems rely on the causal, overlap, and oracle assumptions listed above.

assumptions (6)
  • domain assumption Assumption 3.1 (SCAR): P(D=1,O=d|X) = P(D=1|X)P(O=d|D=1)
    Enables estimation of the treatment propensity from labeled positives; without it, point identification of the censoring propensity is lost.
  • domain assumption Assumption 3.2 (Unconfoundedness): (Y(1),Y(0)) independent of (O,D) given X
    Standard causal assumption needed to identify E[Y(1)] and E[Y(0)] from observed groups.
  • domain assumption Assumption 3.3 (Common support): pi0(d|x), g0(d|x), zeta_T0(x), zeta_0(x) > c_hold
    Required for inverse-weighting terms to be well-defined and for the efficiency bound to be finite.
  • ad hoc to paper Assumption 4.5 / D.5: propensity score g0 (or e0 and density ratio r0) is known
    Load-bearing oracle assumption for asymptotic normality and efficiency; not satisfied in most applications without auxiliary data.
  • domain assumption Assumption 4.6 (product rate conditions for nuisance estimators)
    Requires cross-fitted nuisance estimators to satisfy product of L2 errors o_p(n^-1/2), standard in double machine learning.
  • domain assumption Consistency/SUTVA encoded in DGP: Y = 1[O=1]Y(1) + 1[O=0](1[D=1]Y(1)+1[D=0]Y(0))
    Defines the observed outcome as the potential outcome under the actual treatment; embedded in the DGP rather than stated as a separate assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of PUATE: Efficient Average Treatment Effect Estimation from Treated (Positive) and Unlabeled Units." pith.science (2026). https://pith.science/paper/XAMYEQEJ

@misc{pith2026250119345,
  author       = {Pith},
  title        = {Pith review of: PUATE: Efficient Average Treatment Effect Estimation from Treated (Positive) and Unlabeled Units},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XAMYEQEJ}},
  note         = {Machine review of arXiv:2501.19345}
}
read the original abstract

The estimation of average treatment effects (ATEs), defined as the difference in expected outcomes between treatment and control groups, is a central topic in causal inference. This study develops semiparametric efficient estimators for ATE in a setting where only a treatment group and an unlabeled group, consisting of units whose treatment status is unknown, are observed. This scenario constitutes a variant of learning from positive and unlabeled data (PU learning) and can be viewed as a special case of ATE estimation with missing data. For this setting, we derive the semiparametric efficiency bounds, which characterize the lowest achievable asymptotic variance for regular estimators. We then construct semiparametric efficient ATE estimators that attain these bounds. Our results contribute to the literature on causal inference with missing data and weakly supervised learning.

Figures

Figures reproduced from arXiv: 2501.19345 by the authors.

Figure 1
Figure 1. Illustration of the censoring and case-control settings [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Empirical distributions of ATE estimates. [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. Empirical distributions of ATE estimates. [PITH_FULL_IMAGE:figures/full_fig_p050_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Empirical distributions of ATE estimates. [PITH_FULL_IMAGE:figures/full_fig_p050_4.png]
Figure 5
Figure 5. Figure 5: Response surface A. Left: censoring setting; Right: case-control setting. [PITH_FULL_IMAGE:figures/full_fig_p051_5.png]
Figure 6
Figure 6. Figure 6: Response surface B. Left: censoring setting; Right: case-control setting. [PITH_FULL_IMAGE:figures/full_fig_p051_6.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Positive-Unlabeled Learning for Control Group Construction in Observational Causal Inference

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Using only treated units and an unlabeled pool, positive-unlabeled learning can select control groups whose ATE estimates match the true effect in simulations and two agricultural case studies.

Reference graph

Works this paper leans on

77 extracted references · 46 canonical work pages · cited by 1 Pith paper

  1. [1]

    Gruber, and Samiran Sinha

    Jaeil Ahn, Bhramar Mukherjee, Stephen B. Gruber, and Samiran Sinha. Missing exposure data in stereotype regression model: Application to matched case--control study with disease subclassification. Biometrics, 67 0 (2): 0 546--558, 2011

  2. [2]

    Alaa and Mihaela van der Schaar

    Ahmed M. Alaa and Mihaela van der Schaar. Bayesian inference of individualized treatment effects using multi-task gaussian processes. In Conference on Neural Information Processing Systems (NeurIPS), pp.\ 3427--3435. Curran Associates Inc., 2017

  3. [3]

    Heejung Bang and James M. Robins. Doubly robust estimation in missing data and causal inference models. Biometrics, 61 0 (4): 0 962--973, 2005

  4. [4]

    Learning from positive and unlabeled data under the selected at random assumption

    Jessa Bekker and Jesse Davis. Learning from positive and unlabeled data under the selected at random assumption. In Proceedings of the Second International Workshop on Learning with Imbalanced Domains: Theory and Applications, 2018

  5. [5]

    Learning from positive and unlabeled data: a survey

    Jessa Bekker and Jesse Davis. Learning from positive and unlabeled data: a survey. Machine Learning, 109 0 (4): 0 719--760, 2020

  6. [6]

    Black, Paul J

    Sandra E. Black, Paul J. Devereux, and Kjell G. Salvanes. Staying in the classroom and out of the maternity ward? the effect of compulsory schooling laws on teenage births. The Economic Journal, 118 0 (530): 0 1025--1054, 2008

  7. [7]

    A general framework for treatment effect estimation in semi-supervised and high dimensional settings, 2024

    Abhishek Chakrabortty and Guorong Dai. A general framework for treatment effect estimation in semi-supervised and high dimensional settings, 2024. a rXiv:2201.00468

  8. [8]

    Double/debiased machine learning for treatment and structural parameters

    Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal, 2018

Show all 77 references
  1. [9]

    Nonparametric estimation of heterogeneous treatment effects: From theory to learning algorithms

    Alicia Curth and Mihaela van der Schaar. Nonparametric estimation of heterogeneous treatment effects: From theory to learning algorithms. In Proceedings of the 24th International Conference on Artificial Intelligence and Statistics (AISTATS), 2021

  2. [10]

    Sugiyama

    Marthinus Christoffel du Plessis and M. Sugiyama. Class prior estimation from positive and unlabeled data. IEICE Transactions on Information and Systems , E97-D 0 (5): 0 1358--1362, 2014

  3. [11]

    Analysis of learning from positive and unlabeled data

    Marthinus Christoffel du Plessis, Gang Niu, and Masashi Sugiyama. Analysis of learning from positive and unlabeled data. In Advances in Neural Information Processing Systems (NeurIPS), pp.\ 703--711, 2014

  4. [12]

    Niu, and Masashi Sugiyama

    Marthinus Christoffel du Plessis, Gang. Niu, and Masashi Sugiyama. Convex formulation for learning from positive and unlabeled data. In International Conference on Machine Learning (ICML), 2015

  5. [13]

    Learning classifiers from only positive and unlabeled data

    Charles Elkan and Keith Noto. Learning classifiers from only positive and unlabeled data. In International Conference on Knowledge Discovery and Data Mining (KDD), 2008

  6. [14]

    On the role of the propensity score in efficient semiparametric estimation of average treatment effects

    Jinyong Hahn. On the role of the propensity score in efficient semiparametric estimation of average treatment effects. Econometrica, 66 0 (2): 0 315--331, 1998

  7. [15]

    Learning disentangled representations for counterfactual regression

    Negar Hassanpour and Russell Greiner. Learning disentangled representations for counterfactual regression. In International Conference on Learning Representations, 2020

  8. [16]

    Mismeasured variables in econometric analysis: Problems from the right and problems from the left

    Jerry Hausman. Mismeasured variables in econometric analysis: Problems from the right and problems from the left. Journal of Economic Perspectives, 15 0 (4), 2001

  9. [17]

    Heckman, Hidehiko Ichimura, and Petra E

    James J. Heckman, Hidehiko Ichimura, and Petra E. Todd. Matching as an econometric evaluation estimator: Evidence from evaluating a job training programme. The Review of Economic Studies, 64 0 (4): 0 605--654, 1997

  10. [18]

    A paradox concerning nuisance parameters and projected estimating functions

    Masayuki Henmi and Shinto Eguchi. A paradox concerning nuisance parameters and projected estimating functions . Biometrika, 2004

  11. [19]

    Jennifer L. Hill. Bayesian nonparametric modeling for causal inference. Journal of Computational and Graphical Statistics, 20 0 (1): 0 217--240, 2011

  12. [20]

    Efficient estimation of average treatment effects using the estimated propensity score

    Keisuke Hirano, Guido Imbens, and Geert Ridder. Efficient estimation of average treatment effects using the estimated propensity score. Econometrica, 2003

  13. [21]

    Horvitz and Donovan J

    Daniel G. Horvitz and Donovan J. Thompson. A generalization of sampling without replacement from a finite universe. Journal of the American Statistical Association, 47 0 (260): 0 663--685, 1952

  14. [22]

    Classification from positive, unlabeled and biased negative data

    Yu-Guan Hsieh, Gang Niu, and Masashi Sugiyama. Classification from positive, unlabeled and biased negative data. In International Conference on Machine Learning (ICML), 2019

  15. [23]

    Imbens and Tony Lancaster

    Guido W. Imbens and Tony Lancaster. Efficient estimation and stratified sampling. Journal of Econometrics, 74 0 (2): 0 289--318, 1996

  16. [24]

    Imbens and Donald B

    Guido W. Imbens and Donald B. Rubin. Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. Cambridge University Press, 2015

  17. [25]

    Imbens and Jeffrey M

    Guido W. Imbens and Jeffrey M. Wooldridge. Recent developments in the econometrics of program evaluation. Journal of Economic Literature, 47 0 (1): 0 5--86, 2009

  18. [26]

    Johansson, Uri Shalit, and David Sontag

    Fredrik D. Johansson, Uri Shalit, and David Sontag. Learning representations for counterfactual inference. In International Conference on Machine Learning, pp.\ 3020--3029, 2016

  19. [27]

    A least-squares approach to direct importance estimation

    Takafumi Kanamori, Shohei Hido, and Masashi Sugiyama. A least-squares approach to direct importance estimation. Journal of Machine Learning Research, 10 0 (48): 0 1391--1445, 2009

  20. [28]

    Non-negative bregman divergence minimization for deep direct density ratio estimation

    Masahiro Kato and Takeshi Teshima. Non-negative bregman divergence minimization for deep direct density ratio estimation. In International Conference on Machine Learning (ICML), 2021

  21. [29]

    Alternate estimation of a classifier and the class-prior from positive and unlabeled data, 2018

    Masahiro Kato, Liyuan Xu, Gang Niu, and Masashi Sugiyama. Alternate estimation of a classifier and the class-prior from positive and unlabeled data, 2018. a rXiv:1809.05710

  22. [30]

    Learning from positive and unlabeled data with a selection bias

    Masahiro Kato, Takeshi Teshima, and Junya Honda. Learning from positive and unlabeled data with a selection bias. In International Conference on Learning Representations (ICLR), 2019

  23. [31]

    Efficient adaptive experimental design for average treatment effect estimation, 2020

    Masahiro Kato, Takuya Ishihara, Junya Honda, and Yusuke Narita. Efficient adaptive experimental design for average treatment effect estimation, 2020. a rXiv:2002.05308

  24. [32]

    The adaptive doubly robust estimator and a paradox concerning logging policy

    Masahiro Kato, Kenichiro McAlinn, and Shota Yasui. The adaptive doubly robust estimator and a paradox concerning logging policy. In International Conference on Neural Information Processing Systems (NeurIPS), 2021

  25. [33]

    Active adaptive experimental design for treatment effect estimation with covariate choice

    Masahiro Kato, Akihiro Oga, Wataru Komatsubara, and Ryo Inokuchi. Active adaptive experimental design for treatment effect estimation with covariate choice. In International Conference on Machine Learning (ICML), 2024

  26. [34]

    Edward H. Kennedy. Semiparametric theory and empirical processes in causal inference, 2016. a rXiv: 1510.04740

  27. [35]

    Edward H. Kennedy. Efficient nonparametric causal inference with missing exposure information. The International Journal of Biostatistics, 16 0 (1), 2020

  28. [36]

    Edward H. Kennedy. Semiparametric doubly robust targeted double machine learning: a review, 2023. a rXiv: 2203.06469

  29. [37]

    Kennedy, Sivaraman Balakrishnan, James M

    Edward H. Kennedy, Sivaraman Balakrishnan, James M. Robins, and Larry Wasserman. Minimax rates for heterogeneous causal effect estimation. The Annals of Statistics, 52 0 (2): 0 793 -- 816, 2024

  30. [38]

    Positive-unlabeled learning with non-negative risk estimator

    Ryuichi Kiryo, Gang Niu, Marthinus Christoffel du Plessis, and Masashi Sugiyama. Positive-unlabeled learning with non-negative risk estimator. In Advances in Neural Information Processing Systems (NeurIPS), 2017

  31. [39]

    Chris A. J. Klaassen. Consistent estimation of the influence function of locally asymptotically linear estimators. Annals of Statistics, 15, 1987

  32. [40]

    Estimating conditional average treatment effects with missing treatment information

    Milan Kuzmanovic, Tobias Hatt, and Stefan Feuerriegel. Estimating conditional average treatment effects with missing treatment information. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2023

  33. [41]

    Case-control studies with contaminated controls

    Tony Lancaster and Guido Imbens. Case-control studies with contaminated controls. Journal of Econometrics, 71 0 (1): 0 145--160, 1996

  34. [42]

    Estimation of average treatment effects with misclassification

    Arthur Lewbel. Estimation of average treatment effects with misclassification. Econometrica, 75 0 (2): 0 537--551, 2007

  35. [43]

    Roderick Little and Donald B. Rubin. Statistical analysis with missing data. Wiley series in probability and mathematical statistics. Probability and mathematical statistics. Wiley, 2002

  36. [44]

    Identification and estimation of regression models with misclassification

    Aprajit Mahajan. Identification and estimation of regression models with misclassification. Econometrica, 74 0 (3): 0 631--665, 2006

  37. [45]

    Charles F. Manski. Identification problems in the social sciences. Sociological Methodology, 23: 0 1--56, 1993

  38. [46]

    Charles F. Manski. Partial Identification in Econometrics, pp.\ 178--188. Palgrave Macmillan UK, 2010

  39. [47]

    Missing treatments

    Francesca Molinari. Missing treatments. Journal of Business & Economic Statistics, 28 0 (1): 0 82--95, 2010

  40. [48]

    Sur les applications de la theorie des probabilites aux experiences agricoles: Essai des principes

    Jerzy Neyman. Sur les applications de la theorie des probabilites aux experiences agricoles: Essai des principes. Statistical Science, 5: 0 463--472, 1923

  41. [49]

    Nie and S

    X. Nie and S. Wager. Quasi-oracle estimation of heterogeneous treatment effects. Biometrika, 108, 2020

  42. [50]

    Theoretical comparisons of positive-unlabeled learning against positive-negative learning

    Gang Niu, Marthinus Christoffel du Plessis, Tomoya Sakai, Yao Ma, and Masashi Sugiyama. Theoretical comparisons of positive-unlabeled learning against positive-negative learning. In Advances in Neural Information Processing Systems (NeurIPS), 2016

  43. [51]

    Mixture proportion estimation via kernel embeddings of distributions

    Harish Ramaswamy, Clayton Scott, and Ambuj Tewari. Mixture proportion estimation via kernel embeddings of distributions. In International Conference on Machine Learning (ICML), 2016

  44. [52]

    Donald B. Rubin. Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66: 0 688--701, 1974

  45. [53]

    Donald B. Rubin. Inference and missing data. Biometrika, 63 0 (3): 0 581--592, 1976

  46. [54]

    Semi-supervised classification based on classification from positive and unlabeled data

    Tomoya Sakai, Marthinus Christoffel du Plessis, Gang Niu, and Masashi Sugiyama. Semi-supervised classification based on classification from positive and unlabeled data. In International Conference on Machine Learning (ICML), 2017

  47. [55]

    Nonparametric regression using deep neural networks with relu activation function

    Johannes Schmidt-Hieber. Nonparametric regression using deep neural networks with relu activation function. The Annals of Statistics, 48 0 (4), 2020

  48. [56]

    Introduction to modern causal inference, 2024

    Alejandro Schuler and Mark van der Laan. Introduction to modern causal inference, 2024. URL https://alejandroschuler.github.io/mci/introduction-to-modern-causal-inference.html

  49. [57]

    Johansson, and David Sontag

    Uri Shalit, Fredrik D. Johansson, and David Sontag. Estimating individual treatment effect: Generalization bounds and algorithms. In International Conference on Machine Learning (ICML), pp.\ 3076--3085, 2017

  50. [58]

    Blei, and Victor Veitch

    Claudia Shi, David M. Blei, and Victor Veitch. Adapting neural networks for the estimation of treatment effects. In International Conference on Neural Information Processing Systems. Curran Associates Inc., 2019

  51. [59]

    Scott Cardell

    Dan Steinberg and N. Scott Cardell. Estimating logistic regression models when the dependent variable has no variance. Communications in Statistics - Theory and Methods, 21 0 (2): 0 423--450, 1992

  52. [60]

    Direct importance estimation for covariate shift adaptation

    Masashi Sugiyama, Taiji Suzuki, Shinichi Nakajima, Hisashi Kashima, Paul von B \"u nau, and Motoaki Kawanabe. Direct importance estimation for covariate shift adaptation. Annals of the Institute of Statistical Mathematics, 60 0 (4): 0 699--746, 2008

  53. [61]

    Density Ratio Estimation in Machine Learning

    Masashi Sugiyama, Taiji Suzuki, and Takafumi Kanamori. Density Ratio Estimation in Machine Learning. Cambridge University Press, 2012

  54. [62]

    Machine Learning from Weak Supervision: An Empirical Risk Minimization Approach (Adaptive Computation and Machine Learning series)

    Masashi Sugiyama, Han Bao, Takashi Ishida, Nan Lu, and Tomoya Sakai. Machine Learning from Weak Supervision: An Empirical Risk Minimization Approach (Adaptive Computation and Machine Learning series). The MIT Press, 2022

  55. [63]

    Learning from biased positive-unlabeled data via threshold calibration

    Pawe Teisseyre, Timo Martens, Jessa Bekker, and Jesse Davis. Learning from biased positive-unlabeled data via threshold calibration. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2025

  56. [64]

    A. Tsiatis. Semiparametric Theory and Missing Data. Springer Series in Statistics. Springer New York, 2007

  57. [65]

    Tsybakov

    Alexandre B. Tsybakov. Introduction to Nonparametric Estimation. Springer Publishing Company, Incorporated, 1st edition, 2008

  58. [66]

    Off-policy evaluation and learning for external validity under a covariate shift

    Masatoshi Uehara, Masahiro Kato, and Shota Yasui. Off-policy evaluation and learning for external validity under a covariate shift. In Conference on Neural Information Processing Systems (NeurIPS), 2020

  59. [67]

    Targeted maximum likelihood learning, 2006

    van der Laan. Targeted maximum likelihood learning, 2006. U.C. Berkeley Division of Biostatistics Working Paper Series. Working Paper 213. https://biostats.bepress.com/ucbbiostat/paper213/

  60. [68]

    van der Vaart

    Aad W. van der Vaart. Asymptotic Statistics. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press, 1998

  61. [69]

    van der Vaart

    Aad W. van der Vaart. Semiparametric statistics, 2002. URL https://sites.stat.washington.edu/jaw/COURSES/EPWG/stflour.pdf

  62. [70]

    Estimation and inference of heterogeneous treatment effects using random forests

    Stefan Wager and Susan Athey. Estimation and inference of heterogeneous treatment effects using random forests. Journal of the American Statistical Association, 113 0 (523): 0 1228--1242, 2018

  63. [71]

    Wooldridge

    Jeffrey M. Wooldridge. Asymptotic properties of weighted m-estimation for standard stratified samples. Econometric Theory, 2001

  64. [72]

    Uplift modeling from separate labels

    Ikko Yamane, Florian Yger, Jamal Atif, and Masashi Sugiyama. Uplift modeling from separate labels. In International Conference on Neural Information Processing Systems (NeurIPS), pp.\ 9949--9959. Curran Associates Inc., 2018

  65. [73]

    Doubly robust calibration of prediction sets under covariate shift

    Yachong Yang, Arun Kumar Kuchibhotla, and Eric Tchetgen Tchetgen. Doubly robust calibration of prediction sets under covariate shift. Journal of the Royal Statistical Society Series B: Statistical Methodology, 86 0 (4): 0 943--965, 03 2024

  66. [74]

    Causal inference with missing exposure information: Methods and applications to an obstetric study

    Zhiwei Zhang, Wei Liu, Bo Zhang, Li Tang, and Jun Zhang. Causal inference with missing exposure information: Methods and applications to an obstetric study. Statistical methods in medical research, 25: 0 1003--1014, 12 2013

  67. [75]

    To adjust or not to adjust? estimating the average treatment effect in randomized experiments with missing covariates

    Anqi Zhao and Peng Ding. To adjust or not to adjust? estimating the average treatment effect in randomized experiments with missing covariates. Journal of the American Statistical Association, 119 0 (545): 0 450--460, 2024

  68. [76]

    Cross-validated targeted minimum-loss-based estimation

    Wenjing Zheng and Mark J van der Laan. Cross-validated targeted minimum-loss-based estimation. In Targeted Learning: Causal Inference for Observational and Experimental Data. Springer New York, NY, 2011

  69. [77]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.