REVIEW 3 major objections 4 minor 1 cited by
PUATE: Efficient Average Treatment Effect Estimation from Treated (Positive) and Unlabeled Units
T0 review · 3 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read This paper claims that average treatment effect (ATE) estimation from treated-positive and unlabeled data has an attainable semiparametric efficiency bound, and constructs cross-fitted estimators that attain it in both censoring and…
desk verdict Careful EIF derivation for PU-style ATE, but the headline efficiency claim only holds with known censoring propensity; the abstract overstates it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the efficient influence function (EIF), the score whose variance defines the efficiency bound and whose sample average, with estimated nuisance parameters, gives the optimal estimator. In the censoring setting it is $\Psi_{\mathrm{cens}}=S_{\mathrm{cens}}-\tau_0$, where $S_{\mathrm{cens}}$ combines inverse-probability-weighted residuals of the observed treated outcomes and of the unlabeled-pool outcome, plus a level correction involving the conditional means $\mu_{T,0}$ and $\nu_0$; the censoring propensity score $g_0(d\mid X)=P(D=d\mid X,O=0)$ enters as a known weight. The estimator is the sample average of $S_{\mathrm{cens}}$ with cross-fitted nuisance estimates, so the EIF acts as the estimating equation. For the case-control setting, the corresponding objects are the pair of influence functions $S_{cc}^{(T)}$ and $S_{cc}^{(U)}$, where the case-control propensity $e_0$ and the covariate density ratio $r_0=\zeta_0/\zeta_{T,0}$ are the oracle inputs.
What would settle it
Simulate a DGP satisfying Assumptions 3.2, 3.3, 4.5, and 4.6 with the censoring propensity score known, run the cross-fitted estimator at $n=3000$, and compare the empirical variance of $\sqrt{n}(\hat{\tau}_n^{\mathrm{cens\text{-}eff}}-\tau_0)$ with $V_{\mathrm{cens}}$; Theorem 4.7 fails if the discrepancy exceeds Monte Carlo error. A complementary check is to estimate $g_0$ from the same primary sample without an auxiliary dataset; if $\sqrt{n}$-normality at variance $V_{\mathrm{cens}}$ still holds, the paper's claim that the oracle $g_0$ is needed is contradicted.
Extended reading notes
Core claim
The paper's central claim is that the ATE in the PU setting has an attainable semiparametric efficiency bound. In the censoring setting, the efficient influence function is $\Psi_{\mathrm{cens}}(X,O,Y;\mu_{T,0},\nu_0,\pi_0,g_0)=S_{\mathrm{cens}}-\tau_0$, and the efficiency bound is $V_{\mathrm{cens}}=\mathbb{E}[\Psi_{\mathrm{cens}}^2]$; Theorem 4.7 shows that a cross-fitted estimator formed by averaging $S_{\mathrm{cens}}$ with estimated outcome regressions and observation probabilities, but with the censoring propensity score $g_0$ known, satisfies $\sqrt{n}(\hat{\tau}_n^{\mathrm{cens\text{-}eff}}-\tau_0)\xrightarrow{d}N(0,V_{\mathrm{cens}})$. In the case-control setting, Theorem D.7 gives the analogous result with the case-control propensity score $e_0$ and density ratio $r_0$ known. The paper further shows that when $g_0$ is estimated from the same data without an auxiliary sample, it cannot establish $\sqrt{n}$-consistency; with an auxiliary covariate-only sample, Corollary 4.9 restores asymptotic normality.
Load-bearing premise
The load-bearing premise is that we know the probability that an unlabeled unit actually received the treatment, given its covariates, or can estimate that probability from a separate auxiliary dataset; if that probability is estimated only from the primary data, the paper's efficiency claim does not go through.
Editorial extensions
If this is right
- If Theorem 4.7 holds, the proposed censoring-setting estimator is asymptotically optimal: its variance equals the efficiency bound, so no regular estimator using the same information can have smaller asymptotic variance.
- If Theorem D.7 holds, the case-control estimator attains its bound using only consistent outcome regressions once $e_0$ and $r_0$ are known; the nuisance estimation burden is lighter than in the censoring setting.
- With consistent $g_0$, consistency of the censoring estimator is doubly robust: it survives if either the observation model $\pi_0$ or both outcome regressions $\mu_{T,0}$ and $\nu_0$ are consistent.
- The IPW estimator has asymptotic variance at least $V_{\mathrm{cens}}$, so the efficient estimator strictly improves on it whenever the outcome means are nonzero.
- When $g_0$ is estimated from the same primary data without an auxiliary sample, the paper's asymptotic-normality guarantee does not apply; practitioners are left with consistency, not efficiency.
Reading between the lines
- An implication the paper leaves implicit is that the main practical obstacle is assembling the auxiliary covariate-only data needed to estimate $g_0$; the paper does not analyze how large $n_{\mathrm{aux}}$ must be relative to $n$ to make the efficiency approximation tight.
- A natural testable extension is to derive the exact asymptotic variance when $g_0$ is estimated from an auxiliary sample of size $n_{\mathrm{aux}}$ and to show how the efficiency bound is approached as $n_{\mathrm{aux}}\to\infty$.
- The censoring setting's oracle assumption suggests a partial-identification reading: without a reliable $g_0$, the ATE may still be bounded rather than point-estimated, and connecting this to existing bounds for missing-treatment data would be a direct follow-up.
- In the case-control setting, the density ratio $r_0$ is put on the same oracle footing as $e_0$; since direct density-ratio estimation is often harder than propensity estimation, whether plug-in $r_0$ estimators preserve efficiency is an open question this paper does not settle.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies estimation of the average treatment effect when the data consist of a treated (positive) group and an unlabeled group whose treatment status is unknown, a setting the authors call PUATE. Two sampling schemes are considered: a censoring setting with a single dataset containing observed treated units and units with missing treatment labels, and a case-control setting with separate treated and unlabeled samples. The authors derive efficient influence functions and semiparametric efficiency bounds for both settings, propose cross-fitted estimators based on those influence functions, and prove consistency and asymptotic normality under rate conditions on nuisance estimators. The central theoretical results are Theorem 4.7 for the censoring setting and Theorem D.7 for the case-control setting. The paper also reports simulations and semi-synthetic IHDP experiments.
Significance. If the results are read as an oracle-efficiency theory, the paper makes a useful contribution: it explicitly derives the efficient influence functions for a PU-type missing-treatment problem, constructs cross-fitted estimators, and provides the corresponding asymptotic normality results. The paper is also transparent about its main limitation, namely that asymptotic normality requires the censoring propensity score (or the case-control propensity score and density ratio) to be known or estimated from an independent auxiliary dataset. I found no circularity in the derivations; the issue is instead one of identification relative to the model stated in the abstract. Once the claims are re-framed as oracle-model efficiency results, the technical machinery is informative and the simulation study, though limited, is a reasonable check on the finite-sample behavior of the proposed estimators. The uncorrected efficiency bound in Theorem 4.2 and the overstatement of the efficiency claim in the abstract are the main obstacles to publication in the current form.
major comments (3)
- [Abstract; §4.1 (Lemma 4.1, Theorem 4.2); Appendix H; §4.4-4.5] The efficiency claim is stated for the model described by Assumptions 3.2–3.3, but Appendix H derives the efficient influence function in parametric submodels that keep g0(d|X) fixed: the submodel density is p_etilde(y|x;θ) = g0(1|x)p_Y(1)(y|x;θ) + g0(0|x)p_Y(0)(y|x;θ). Therefore the tangent space contains no scores for the censoring propensity score g0, and V_cens is the efficiency bound for an oracle model in which g0 is known, not for the PUATE model in which g0 is an unknown component of the data-generating process. In the latter model τ0 is not identified: for any observed law of (X,O,Y), one can change g0(1|X) and redefine p_Y(0) so that the mixture q(y|x) = g0(1|x)p_Y(1)(y|x) + (1-g0(1|x))p_Y(0)(y|x) is unchanged while E[Y(0)] changes. Consequently no regular √n estimator exists in that model, and the abstract's sentence 'We then construct semiparametric efficient ATE estimators that attain these bounds' is not supported. Theorem 4.7 explicitly requires Assumption 4.5 (known g0), and Section 4.5 concedes that √n-consistency cannot be obtained when g0 is estimated from the same data. The same issue applies to the case-control setting, where Assumption D.5 fixes e0 and r0. I recommend re-stating Lemma 4.1, Theorem 4.2, and the case-control analogues as oracle-model results, or as results with an independent auxiliary dataset (Corollary 4.9), and revising the abstract accordingly.
- [§4.1, Theorem 4.2] The displayed variance formula in Theorem 4.2 is inconsistent with the efficient influence function in Lemma 4.1. Writing S_cens for the influence-function score, the O=1 terms combine with coefficient 1 + g0(1|X)/g0(0|X) = 1/g0(0|X), not (1 - g0(1|X)/g0(0|X)), and the O=0 part has variance Var(eY|X)/π0(0|X), not Var(eY|X)/π0(1|X). The correct expression should be E[ (1/g0(0|X)^2)(Var(Y(1)|X)/π0(1|X) + Var(eY|X)/π0(0|X)) + (τ0(X)-τ0)^2 ]. This error is load-bearing because V_cens is the claimed efficiency bound used in Theorem 4.7, Corollary 4.9, and the simulation discussion. Please correct the theorem and verify all uses of V_cens.
- [§5, Theorem 5.1; Appendix D, Theorem D.7] The case-control efficiency claim has the same oracle-conditional structure as the censoring result. Theorem 5.1 and Theorem D.7 assume that e0(d|X) and r0(X) are known (Assumption D.5). Without such knowledge, or without additional identification assumptions such as an identifiable class prior, the unlabeled sample is a mixture with unknown weight e0(1|X), and τ0 is not identified from the stated data-generating process. The abstract presents the case-control result as an unconditional efficiency bound, which is not supported by the formal theorems. Please make the oracle assumption explicit in the main text and abstract for both settings, and state precisely which quantities are assumed known versus estimated.
minor comments (4)
- [Appendix I.2, proof of Theorem 4.7] In the displayed calculation of the second-order remainder, the third term in the expression for the conditional expectation appears with bπ_n(0|X) in the denominator where the EIF has bπ_n(1|X), and similarly π0(0) appears where π0(1) is required. This appears to be a typo, but it should be corrected because it is part of the proof of the main asymptotic normality result.
- [§2.3] The definition of Yi uses eY_i before eY_i is defined; reorder the two displayed equations for readability.
- [§4.1 and §4.2] The phrases 'If Assumptions 3.2–3.3' and 'such a derivation of the efficient estimator as the estimation equation approach' are incomplete; please edit for grammatical completeness.
- [Appendix M, Tables 4-5] In the semi-synthetic IHDP censoring experiments, the proposed efficient estimator has coverage ratios of 0.22 (response surface A) and 0.01 (response surface B), far below the nominal level. The paper should comment on this discrepancy and whether the asymptotic approximation is expected to be reliable at n = 747 with estimated g0.
Circularity Check
No circular reasoning identified; the oracle propensity-score assumption is a stated limitation rather than a circular step.
full rationale
The paper's central result—semiparametric efficiency of the proposed ATE estimators in the censoring and case-control PU settings—is derived from standard semiparametric theory (van der Vaart 1998; Hahn 1998) by computing the efficient influence function, not by assuming the estimator's own variance or the target value. The formulas for Ψcens and the bound Vcens are obtained from the score of parametric submodels and the Riesz representation theorem, and the asymptotic normality in Theorem 4.7 is a standard cross-fitting argument. The main caveat, which the paper itself states in Assumption 4.5 and Section 4.5, is that the efficiency and √n-normality results require the censoring propensity score g0 to be known (or estimated from an auxiliary dataset), and the proof of Lemma 4.1 fixes g0 in the parametric submodel. This is an oracle-model limitation and a possible mismatch with the abstract's unconditional wording, but it is not circular: the bound and the estimator are not defined in terms of each other, and the regularity conditions on nuisance estimators are stated independently. Citations to Uehara et al. (2020) and Kato et al. (2024) provide the general stratified-sampling efficiency framework; they do not assume the PUATE efficiency result, so the self-citations are technical and not load-bearing. No step reduces by construction to its own input.
Assumptions & free parameters
assumptions (6)
- domain assumption Assumption 3.1 (SCAR): P(D=1,O=d|X) = P(D=1|X)P(O=d|D=1)
- domain assumption Assumption 3.2 (Unconfoundedness): (Y(1),Y(0)) independent of (O,D) given X
- domain assumption Assumption 3.3 (Common support): pi0(d|x), g0(d|x), zeta_T0(x), zeta_0(x) > c_hold
- ad hoc to paper Assumption 4.5 / D.5: propensity score g0 (or e0 and density ratio r0) is known
- domain assumption Assumption 4.6 (product rate conditions for nuisance estimators)
- domain assumption Consistency/SUTVA encoded in DGP: Y = 1[O=1]Y(1) + 1[O=0](1[D=1]Y(1)+1[D=0]Y(0))
Cite this review
Pith. "Pith review of PUATE: Efficient Average Treatment Effect Estimation from Treated (Positive) and Unlabeled Units." pith.science (2026). https://pith.science/paper/XAMYEQEJ
@misc{pith2026250119345,
author = {Pith},
title = {Pith review of: PUATE: Efficient Average Treatment Effect Estimation from Treated (Positive) and Unlabeled Units},
year = {2026},
howpublished = {\url{https://pith.science/paper/XAMYEQEJ}},
note = {Machine review of arXiv:2501.19345}
}
read the original abstract
The estimation of average treatment effects (ATEs), defined as the difference in expected outcomes between treatment and control groups, is a central topic in causal inference. This study develops semiparametric efficient estimators for ATE in a setting where only a treatment group and an unlabeled group, consisting of units whose treatment status is unknown, are observed. This scenario constitutes a variant of learning from positive and unlabeled data (PU learning) and can be viewed as a special case of ATE estimation with missing data. For this setting, we derive the semiparametric efficiency bounds, which characterize the lowest achievable asymptotic variance for regular estimators. We then construct semiparametric efficient ATE estimators that attain these bounds. Our results contribute to the literature on causal inference with missing data and weakly supervised learning.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
Positive-Unlabeled Learning for Control Group Construction in Observational Causal Inference
Using only treated units and an unlabeled pool, positive-unlabeled learning can select control groups whose ATE estimates match the true effect in simulations and two agricultural case studies.
Reference graph
Works this paper leans on
-
[1]
Jaeil Ahn, Bhramar Mukherjee, Stephen B. Gruber, and Samiran Sinha. Missing exposure data in stereotype regression model: Application to matched case--control study with disease subclassification. Biometrics, 67 0 (2): 0 546--558, 2011
work page 2011
-
[2]
Alaa and Mihaela van der Schaar
Ahmed M. Alaa and Mihaela van der Schaar. Bayesian inference of individualized treatment effects using multi-task gaussian processes. In Conference on Neural Information Processing Systems (NeurIPS), pp.\ 3427--3435. Curran Associates Inc., 2017
work page 2017
-
[3]
Heejung Bang and James M. Robins. Doubly robust estimation in missing data and causal inference models. Biometrics, 61 0 (4): 0 962--973, 2005
2005
-
[4]
Learning from positive and unlabeled data under the selected at random assumption
Jessa Bekker and Jesse Davis. Learning from positive and unlabeled data under the selected at random assumption. In Proceedings of the Second International Workshop on Learning with Imbalanced Domains: Theory and Applications, 2018
work page 2018
-
[5]
Learning from positive and unlabeled data: a survey
Jessa Bekker and Jesse Davis. Learning from positive and unlabeled data: a survey. Machine Learning, 109 0 (4): 0 719--760, 2020
work page 2020
-
[6]
Sandra E. Black, Paul J. Devereux, and Kjell G. Salvanes. Staying in the classroom and out of the maternity ward? the effect of compulsory schooling laws on teenage births. The Economic Journal, 118 0 (530): 0 1025--1054, 2008
work page 2008
-
[7]
Abhishek Chakrabortty and Guorong Dai. A general framework for treatment effect estimation in semi-supervised and high dimensional settings, 2024. a rXiv:2201.00468
arXiv 2024
-
[8]
Double/debiased machine learning for treatment and structural parameters
Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal, 2018
2018
Show all 77 references
-
[9]
Nonparametric estimation of heterogeneous treatment effects: From theory to learning algorithms
Alicia Curth and Mihaela van der Schaar. Nonparametric estimation of heterogeneous treatment effects: From theory to learning algorithms. In Proceedings of the 24th International Conference on Artificial Intelligence and Statistics (AISTATS), 2021
2021
-
[10]
Sugiyama
Marthinus Christoffel du Plessis and M. Sugiyama. Class prior estimation from positive and unlabeled data. IEICE Transactions on Information and Systems , E97-D 0 (5): 0 1358--1362, 2014
2014
-
[11]
Analysis of learning from positive and unlabeled data
Marthinus Christoffel du Plessis, Gang Niu, and Masashi Sugiyama. Analysis of learning from positive and unlabeled data. In Advances in Neural Information Processing Systems (NeurIPS), pp.\ 703--711, 2014
2014
-
[12]
Niu, and Masashi Sugiyama
Marthinus Christoffel du Plessis, Gang. Niu, and Masashi Sugiyama. Convex formulation for learning from positive and unlabeled data. In International Conference on Machine Learning (ICML), 2015
2015
-
[13]
Learning classifiers from only positive and unlabeled data
Charles Elkan and Keith Noto. Learning classifiers from only positive and unlabeled data. In International Conference on Knowledge Discovery and Data Mining (KDD), 2008
2008
-
[14]
On the role of the propensity score in efficient semiparametric estimation of average treatment effects
Jinyong Hahn. On the role of the propensity score in efficient semiparametric estimation of average treatment effects. Econometrica, 66 0 (2): 0 315--331, 1998
1998
-
[15]
Learning disentangled representations for counterfactual regression
Negar Hassanpour and Russell Greiner. Learning disentangled representations for counterfactual regression. In International Conference on Learning Representations, 2020
2020
-
[16]
Mismeasured variables in econometric analysis: Problems from the right and problems from the left
Jerry Hausman. Mismeasured variables in econometric analysis: Problems from the right and problems from the left. Journal of Economic Perspectives, 15 0 (4), 2001
2001
-
[17]
Heckman, Hidehiko Ichimura, and Petra E
James J. Heckman, Hidehiko Ichimura, and Petra E. Todd. Matching as an econometric evaluation estimator: Evidence from evaluating a job training programme. The Review of Economic Studies, 64 0 (4): 0 605--654, 1997
1997
-
[18]
A paradox concerning nuisance parameters and projected estimating functions
Masayuki Henmi and Shinto Eguchi. A paradox concerning nuisance parameters and projected estimating functions . Biometrika, 2004
2004
-
[19]
Jennifer L. Hill. Bayesian nonparametric modeling for causal inference. Journal of Computational and Graphical Statistics, 20 0 (1): 0 217--240, 2011
2011
-
[20]
Efficient estimation of average treatment effects using the estimated propensity score
Keisuke Hirano, Guido Imbens, and Geert Ridder. Efficient estimation of average treatment effects using the estimated propensity score. Econometrica, 2003
2003
-
[21]
Horvitz and Donovan J
Daniel G. Horvitz and Donovan J. Thompson. A generalization of sampling without replacement from a finite universe. Journal of the American Statistical Association, 47 0 (260): 0 663--685, 1952
1952
-
[22]
Classification from positive, unlabeled and biased negative data
Yu-Guan Hsieh, Gang Niu, and Masashi Sugiyama. Classification from positive, unlabeled and biased negative data. In International Conference on Machine Learning (ICML), 2019
2019
-
[23]
Imbens and Tony Lancaster
Guido W. Imbens and Tony Lancaster. Efficient estimation and stratified sampling. Journal of Econometrics, 74 0 (2): 0 289--318, 1996
1996
-
[24]
Imbens and Donald B
Guido W. Imbens and Donald B. Rubin. Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. Cambridge University Press, 2015
2015
-
[25]
Imbens and Jeffrey M
Guido W. Imbens and Jeffrey M. Wooldridge. Recent developments in the econometrics of program evaluation. Journal of Economic Literature, 47 0 (1): 0 5--86, 2009
2009
-
[26]
Johansson, Uri Shalit, and David Sontag
Fredrik D. Johansson, Uri Shalit, and David Sontag. Learning representations for counterfactual inference. In International Conference on Machine Learning, pp.\ 3020--3029, 2016
2016
-
[27]
A least-squares approach to direct importance estimation
Takafumi Kanamori, Shohei Hido, and Masashi Sugiyama. A least-squares approach to direct importance estimation. Journal of Machine Learning Research, 10 0 (48): 0 1391--1445, 2009
2009
-
[28]
Non-negative bregman divergence minimization for deep direct density ratio estimation
Masahiro Kato and Takeshi Teshima. Non-negative bregman divergence minimization for deep direct density ratio estimation. In International Conference on Machine Learning (ICML), 2021
2021
-
[29]
Alternate estimation of a classifier and the class-prior from positive and unlabeled data, 2018
Masahiro Kato, Liyuan Xu, Gang Niu, and Masashi Sugiyama. Alternate estimation of a classifier and the class-prior from positive and unlabeled data, 2018. a rXiv:1809.05710
2018 arXiv
-
[30]
Learning from positive and unlabeled data with a selection bias
Masahiro Kato, Takeshi Teshima, and Junya Honda. Learning from positive and unlabeled data with a selection bias. In International Conference on Learning Representations (ICLR), 2019
2019
-
[31]
Efficient adaptive experimental design for average treatment effect estimation, 2020
Masahiro Kato, Takuya Ishihara, Junya Honda, and Yusuke Narita. Efficient adaptive experimental design for average treatment effect estimation, 2020. a rXiv:2002.05308
2020 arXiv
-
[32]
The adaptive doubly robust estimator and a paradox concerning logging policy
Masahiro Kato, Kenichiro McAlinn, and Shota Yasui. The adaptive doubly robust estimator and a paradox concerning logging policy. In International Conference on Neural Information Processing Systems (NeurIPS), 2021
2021
-
[33]
Active adaptive experimental design for treatment effect estimation with covariate choice
Masahiro Kato, Akihiro Oga, Wataru Komatsubara, and Ryo Inokuchi. Active adaptive experimental design for treatment effect estimation with covariate choice. In International Conference on Machine Learning (ICML), 2024
2024
-
[34]
Edward H. Kennedy. Semiparametric theory and empirical processes in causal inference, 2016. a rXiv: 1510.04740
2016 arXiv
-
[35]
Edward H. Kennedy. Efficient nonparametric causal inference with missing exposure information. The International Journal of Biostatistics, 16 0 (1), 2020
2020
-
[36]
Edward H. Kennedy. Semiparametric doubly robust targeted double machine learning: a review, 2023. a rXiv: 2203.06469
2023 arXiv
-
[37]
Kennedy, Sivaraman Balakrishnan, James M
Edward H. Kennedy, Sivaraman Balakrishnan, James M. Robins, and Larry Wasserman. Minimax rates for heterogeneous causal effect estimation. The Annals of Statistics, 52 0 (2): 0 793 -- 816, 2024
2024
-
[38]
Positive-unlabeled learning with non-negative risk estimator
Ryuichi Kiryo, Gang Niu, Marthinus Christoffel du Plessis, and Masashi Sugiyama. Positive-unlabeled learning with non-negative risk estimator. In Advances in Neural Information Processing Systems (NeurIPS), 2017
2017
-
[39]
Chris A. J. Klaassen. Consistent estimation of the influence function of locally asymptotically linear estimators. Annals of Statistics, 15, 1987
1987
-
[40]
Estimating conditional average treatment effects with missing treatment information
Milan Kuzmanovic, Tobias Hatt, and Stefan Feuerriegel. Estimating conditional average treatment effects with missing treatment information. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2023
2023
-
[41]
Case-control studies with contaminated controls
Tony Lancaster and Guido Imbens. Case-control studies with contaminated controls. Journal of Econometrics, 71 0 (1): 0 145--160, 1996
1996
-
[42]
Estimation of average treatment effects with misclassification
Arthur Lewbel. Estimation of average treatment effects with misclassification. Econometrica, 75 0 (2): 0 537--551, 2007
2007
-
[43]
Roderick Little and Donald B. Rubin. Statistical analysis with missing data. Wiley series in probability and mathematical statistics. Probability and mathematical statistics. Wiley, 2002
2002
-
[44]
Identification and estimation of regression models with misclassification
Aprajit Mahajan. Identification and estimation of regression models with misclassification. Econometrica, 74 0 (3): 0 631--665, 2006
2006
-
[45]
Charles F. Manski. Identification problems in the social sciences. Sociological Methodology, 23: 0 1--56, 1993
1993
-
[46]
Charles F. Manski. Partial Identification in Econometrics, pp.\ 178--188. Palgrave Macmillan UK, 2010
2010
-
[47]
Missing treatments
Francesca Molinari. Missing treatments. Journal of Business & Economic Statistics, 28 0 (1): 0 82--95, 2010
2010
-
[48]
Sur les applications de la theorie des probabilites aux experiences agricoles: Essai des principes
Jerzy Neyman. Sur les applications de la theorie des probabilites aux experiences agricoles: Essai des principes. Statistical Science, 5: 0 463--472, 1923
1923
-
[49]
Nie and S
X. Nie and S. Wager. Quasi-oracle estimation of heterogeneous treatment effects. Biometrika, 108, 2020
2020
-
[50]
Theoretical comparisons of positive-unlabeled learning against positive-negative learning
Gang Niu, Marthinus Christoffel du Plessis, Tomoya Sakai, Yao Ma, and Masashi Sugiyama. Theoretical comparisons of positive-unlabeled learning against positive-negative learning. In Advances in Neural Information Processing Systems (NeurIPS), 2016
2016
-
[51]
Mixture proportion estimation via kernel embeddings of distributions
Harish Ramaswamy, Clayton Scott, and Ambuj Tewari. Mixture proportion estimation via kernel embeddings of distributions. In International Conference on Machine Learning (ICML), 2016
2016
-
[52]
Donald B. Rubin. Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66: 0 688--701, 1974
1974
-
[53]
Donald B. Rubin. Inference and missing data. Biometrika, 63 0 (3): 0 581--592, 1976
1976
-
[54]
Semi-supervised classification based on classification from positive and unlabeled data
Tomoya Sakai, Marthinus Christoffel du Plessis, Gang Niu, and Masashi Sugiyama. Semi-supervised classification based on classification from positive and unlabeled data. In International Conference on Machine Learning (ICML), 2017
2017
-
[55]
Nonparametric regression using deep neural networks with relu activation function
Johannes Schmidt-Hieber. Nonparametric regression using deep neural networks with relu activation function. The Annals of Statistics, 48 0 (4), 2020
2020
-
[56]
Introduction to modern causal inference, 2024
Alejandro Schuler and Mark van der Laan. Introduction to modern causal inference, 2024. URL https://alejandroschuler.github.io/mci/introduction-to-modern-causal-inference.html
2024
-
[57]
Johansson, and David Sontag
Uri Shalit, Fredrik D. Johansson, and David Sontag. Estimating individual treatment effect: Generalization bounds and algorithms. In International Conference on Machine Learning (ICML), pp.\ 3076--3085, 2017
2017
-
[58]
Blei, and Victor Veitch
Claudia Shi, David M. Blei, and Victor Veitch. Adapting neural networks for the estimation of treatment effects. In International Conference on Neural Information Processing Systems. Curran Associates Inc., 2019
2019
-
[59]
Scott Cardell
Dan Steinberg and N. Scott Cardell. Estimating logistic regression models when the dependent variable has no variance. Communications in Statistics - Theory and Methods, 21 0 (2): 0 423--450, 1992
1992
-
[60]
Direct importance estimation for covariate shift adaptation
Masashi Sugiyama, Taiji Suzuki, Shinichi Nakajima, Hisashi Kashima, Paul von B \"u nau, and Motoaki Kawanabe. Direct importance estimation for covariate shift adaptation. Annals of the Institute of Statistical Mathematics, 60 0 (4): 0 699--746, 2008
2008
-
[61]
Density Ratio Estimation in Machine Learning
Masashi Sugiyama, Taiji Suzuki, and Takafumi Kanamori. Density Ratio Estimation in Machine Learning. Cambridge University Press, 2012
2012
-
[62]
Machine Learning from Weak Supervision: An Empirical Risk Minimization Approach (Adaptive Computation and Machine Learning series)
Masashi Sugiyama, Han Bao, Takashi Ishida, Nan Lu, and Tomoya Sakai. Machine Learning from Weak Supervision: An Empirical Risk Minimization Approach (Adaptive Computation and Machine Learning series). The MIT Press, 2022
2022
-
[63]
Learning from biased positive-unlabeled data via threshold calibration
Pawe Teisseyre, Timo Martens, Jessa Bekker, and Jesse Davis. Learning from biased positive-unlabeled data via threshold calibration. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2025
2025
-
[64]
A. Tsiatis. Semiparametric Theory and Missing Data. Springer Series in Statistics. Springer New York, 2007
2007
-
[65]
Tsybakov
Alexandre B. Tsybakov. Introduction to Nonparametric Estimation. Springer Publishing Company, Incorporated, 1st edition, 2008
2008
-
[66]
Off-policy evaluation and learning for external validity under a covariate shift
Masatoshi Uehara, Masahiro Kato, and Shota Yasui. Off-policy evaluation and learning for external validity under a covariate shift. In Conference on Neural Information Processing Systems (NeurIPS), 2020
2020
-
[67]
Targeted maximum likelihood learning, 2006
van der Laan. Targeted maximum likelihood learning, 2006. U.C. Berkeley Division of Biostatistics Working Paper Series. Working Paper 213. https://biostats.bepress.com/ucbbiostat/paper213/
2006
-
[68]
van der Vaart
Aad W. van der Vaart. Asymptotic Statistics. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press, 1998
1998
-
[69]
van der Vaart
Aad W. van der Vaart. Semiparametric statistics, 2002. URL https://sites.stat.washington.edu/jaw/COURSES/EPWG/stflour.pdf
2002
-
[70]
Estimation and inference of heterogeneous treatment effects using random forests
Stefan Wager and Susan Athey. Estimation and inference of heterogeneous treatment effects using random forests. Journal of the American Statistical Association, 113 0 (523): 0 1228--1242, 2018
2018
-
[71]
Wooldridge
Jeffrey M. Wooldridge. Asymptotic properties of weighted m-estimation for standard stratified samples. Econometric Theory, 2001
2001
-
[72]
Uplift modeling from separate labels
Ikko Yamane, Florian Yger, Jamal Atif, and Masashi Sugiyama. Uplift modeling from separate labels. In International Conference on Neural Information Processing Systems (NeurIPS), pp.\ 9949--9959. Curran Associates Inc., 2018
2018
-
[73]
Doubly robust calibration of prediction sets under covariate shift
Yachong Yang, Arun Kumar Kuchibhotla, and Eric Tchetgen Tchetgen. Doubly robust calibration of prediction sets under covariate shift. Journal of the Royal Statistical Society Series B: Statistical Methodology, 86 0 (4): 0 943--965, 03 2024
2024
-
[74]
Causal inference with missing exposure information: Methods and applications to an obstetric study
Zhiwei Zhang, Wei Liu, Bo Zhang, Li Tang, and Jun Zhang. Causal inference with missing exposure information: Methods and applications to an obstetric study. Statistical methods in medical research, 25: 0 1003--1014, 12 2013
2013
-
[75]
To adjust or not to adjust? estimating the average treatment effect in randomized experiments with missing covariates
Anqi Zhao and Peng Ding. To adjust or not to adjust? estimating the average treatment effect in randomized experiments with missing covariates. Journal of the American Statistical Association, 119 0 (545): 0 450--460, 2024
2024
-
[76]
Cross-validated targeted minimum-loss-based estimation
Wenjing Zheng and Mark J van der Laan. Cross-validated targeted minimum-loss-based estimation. In Targeted Learning: Causal Inference for Observational and Experimental Data. Springer New York, NY, 2011
2011
-
[77]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.