REVIEW 5 major objections 5 minor 48 references
Graph-attention-based Casual Discovery with Trust Region-navigated Clipping Policy Optimization
T0 review · 5 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper claims that per-edge trust-region clipping makes reinforcement-learning-based causal discovery both steadier and more accurate, and that a scaled dot-product graph attention encoder adds further gains, with best structural…
desk verdict A reasonable incremental RL-for-causal-discovery package with real empirical work, but the central trust-region claim is mathematically unsupported and the paper's own results undercut its headline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the KL-gated clipping rule of TRC, built on the factorization of the graph policy into independent Bernoulli subpolicies, one per directed edge. For each edge, the likelihood ratio is clipped only if the per-edge KL divergence between the new and old policies exceeds a threshold, and otherwise it is retained exactly. The paper argues this gives first-order efficiency with trust-region safety, avoiding both TRPO's cost and PPO's aggregate drift. The supporting encoder, SDGAT, uses scaled dot-product attention in a two-level multi-head design to extract variable features without requiring prior neighborhood information, which the paper claims is better suited to causal discovery than GAT's additive attention.
What would settle it
Run TRC on a 12-node linear-Gaussian dataset and record the joint KL divergence between old and new policies at every update; if the joint KL regularly exceeds the intended trust-region bound while per-edge clipping is active, the central mechanism is not doing its claimed work.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the failure of PPO in causal discovery comes from aggregate constraint deviation: the joint likelihood ratio is the product of $n(n-1)$ per-edge ratios, so even trivial per-edge violations compound. TRC replaces the ratio-triggered clipping of PPO with a trust-region-triggered clipping: each per-edge ratio is clipped to a fixed interval around 1 only when that edge's own KL divergence crosses a threshold, and is left untouched otherwise. This is claimed to keep the joint policy near the old policy while preserving the exploratory freedom of edges that have not drifted. With the SDGAT encoder replacing GAT's additive attention by scaled dot-product attention, TRC-BIC and TRC-BIC2 are reported to obtain the lowest structural Hamming distances on the SynTReN pseudo-real datasets and the CYTO protein-signaling dataset among all compared methods.
Load-bearing premise
The paper assumes, without proof, that gating each edge's update by that edge's own statistical distance from its old policy keeps the whole product-of-edges policy inside a trust region; if this per-edge gate does not control the joint deviation, TRC loses its claimed advantage over PPO.
Editorial extensions
If this is right
- On the 12-node linear-Gaussian and LiNGAM settings, TRC converges to a batch negative reward around -2.35 and stops fluctuating earlier than REINFORCE or PSR, with PPO excluded because its SHD exceeded 35.
- On nonlinear quadratic data, TRC produces graphs with SHD at most 1, effectively recovering the true graph.
- On SynTReN, TRC-BIC and TRC-BIC2 reach SHD 35.0 and 34.9, the best among all compared methods, and on CYTO they reach SHD 9 and 10.
- With TRC-BIC on CYTO, the SDGAT encoder converges to a reward of -5.48 with standard deviation 0.27 and better final metrics than GAT and Transformer encoders.
- The TRC idea is claimed to generalize to other high-dimensional combinatorial optimization problems whose actions decompose into many independent subactions.
Reading between the lines
- Beyond the paper, one can directly test whether the per-edge gating actually controls the joint policy by computing the aggregate KL divergence during TRC training; the paper reports clipping rates but not joint KL values, so this is a concrete way to verify the mechanism.
- If TRC works as claimed, the same per-edge clipping rule could transfer to other combinatorial generators, such as molecule or architecture generation, where the action is a product of many independent choices rather than Bernoulli edges.
- The paper leaves implicit that SDGAT could serve as a general structure-agnostic attention encoder, since its experiments only cover causal discovery and the authors themselves note that transductive and inductive performance remains untested.
- A more principled schedule for the clipping threshold and the trust-region threshold could replace the grid search reported in the appendix, provided a closed-form relation between per-edge KL and joint-policy KL is derived.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes Trust Region-navigated Clipping policy optimization (TRC) for RL-based causal discovery, along with a scaled dot-product graph attention encoder (SDGAT). The authors argue that REINFORCE is prone to local convergence, TRPO is computationally expensive, and PPO suffers from aggregate deviation because the joint likelihood ratio is a product of many per-edge ratios; TRC replaces PPO's ratio trigger by a per-edge KL-divergence trigger. Experiments on synthetic linear-Gaussian, LiNGAM, quadratic, GP, SynTReN, and CYTO datasets compare TRC-BIC/TRC-BIC2, PSR, REINFORCE, and non-RL baselines. The central claims are that TRC is more robust and faster than PPO/REINFORCE/PSR and that SDGAT improves causal encoding.
Significance. The problem is important: RL-based causal discovery needs stable policy optimization for high-dimensional binary action spaces. The paper's identification of aggregate deviation in PPO's product ratio is valid and interesting. The SDGAT encoder is a reasonable extension of GAT, and the experimental campaign is broad, covering both synthetic and real benchmarks. The paper ships explicit pseudocode for all algorithms and attempts to justify every design choice. However, the load-bearing theoretical mechanism for TRC is not established: as defined, it does not bound the joint ratio or the joint KL divergence, so the claimed improvement over PPO is unverified. The contradictory GP result and the tuning of hyperparameters on evaluation datasets further weaken the empirical claims.
major comments (5)
- [V-C, Algorithm 3, Eq. (19)] The central robustness claim is unsupported. Eq. (19) clips the per-edge likelihood ratio only when that edge's own D_KL exceeds sigma. For independent Bernoulli subpolicies, D_KL can be arbitrarily small while the likelihood ratio q/p is large (whenever the old probability p is small), so ratios far outside [1-epsilon, 1+epsilon] are never clipped. Moreover, even if every ratio were clipped, the surrogate in Algorithm 3 line 9 uses the product over all subactions; with n=12 and epsilon=0.2, (1+epsilon)^{n(n-1)} is roughly 3.5e10, so the aggregate deviation the authors attribute to PPO remains present in TRC. No bound is provided on the joint ratio or on D_KL(b, pi_theta | S), and Fig. 7 reports only clipping frequencies, which do not measure trust-region satisfaction. The stated guarantee that TRC 'stays safe in the KL bounds' (Section VI-A) is therefore not established.
- [VI-A, Fig. 7] PPO is excluded from all quantitative comparisons after being described as 'miserable' (SHD over 35). The paper therefore never reports a head-to-head TRC-versus-PPO comparison on the same benchmark; the abstract and contribution claims that TRC outperforms PPO are not supported by any table or figure. At minimum, a table with PPO results and variance across seeds is required before such claims can be evaluated.
- [VI-B, Fig. 9] The GP experiment directly contradicts the paper's general claim. The text states that 'PSR would deliver the best result when combined with BIC, whilst REINFORCE still lag behind the other two RL approaches.' This means on one of the four synthetic settings, the proposed TRC does not outperform PSR. The abstract and introduction claim TRC outperforms 'former RL methods' without qualification. This internal inconsistency must be addressed by either revising the claim or explaining this result.
- [Appendix A, Section VI] The (epsilon, delta) pairs are selected by grid search on the CYTO dataset, and the text says 'the choice of (epsilon-delta) pair for synthetic dataset is obtained likewise,' i.e., tuned on the same datasets later used for final SHD reporting. This selection on the test sets inflates the reported performance and makes comparisons with fixed-default baselines unfair. The paper should use a validation-set split or report the sensitivity of the final results to these hyperparameters.
- [Algorithm 4, Table III] The penalty schedule is internally inconsistent and non-reproducible. Table III lists Lambda_1 = 0 and BIC_u = -1, while the text says lambda_1 is increased with upper bound Lambda_1 and lambda_2 is increased with upper bound BIC_u. With BIC_u = -1, the update lambda_2 <- min(lambda_2 + Delta_2, BIC_u) drives lambda_2 negative, which turns the acyclicity penalty into a reward for cycles; Lambda_1 = 0 prevents lambda_1 from ever becoming positive. In addition, Algorithm 4 requires BIC0, but Table III does not list it. These issues must be corrected and the actual values used in the experiments reported.
minor comments (5)
- [Title and Abstract] The title and abstract use 'Casual Discovery' where 'Causal Discovery' is meant; also 'without priori neighbourhood information' should be 'without a priori neighbourhood information' in the abstract and Section IV-A.
- [Section VI-C, Fig. 10] The caption for Fig. 10 says the CYTO dataset has '14 nodes,' while the text states the graph has 11 nodes and 17 edges; please reconcile this discrepancy.
- [Equation (5)] Equation (5) uses the notation '[X At]_i,j' without defining it; the regression estimate for Xi,j should be spelled out.
- [Section V, Algorithms 1 and 2] The advantage conventions differ between Algorithm 1 line 4 (Rt - Rm - V_omega) and Algorithm 2 line 2 (Rt + Rm - V_omega); this sign inconsistency should be checked and corrected.
- [Section VI-A, Fig. 7 and related text] The claim that TRC 'outperforms former RL methods' is not placed in the context of recent RL-based causal discovery approaches such as CORL or DAG-Actor; the scope of the comparison should be stated.
Circularity Check
No derivation is circular; the mild burden is from tuning (ε, δ) on CYTO and then reporting CYTO as the headline result.
-
fitted input called prediction
[Appendix A (Table II, grid search on CYTO) feeding Section VI-C (CYTO SHD results)]
"As an illustration, the grid search results on the CYTO dataset are shown in TABLE II. They are evaluated by the mean and standard deviation of rewards during training along with the average SHD of the final results. ... TRC-BIC and TRC-BIC2 surpass other methods on SynTReN by reaching the SHD of 35.0 and 34.9 respectively. They also achieve the hitherto best result on the protein dataset with the lowest SHD of 9 and 10 respectively."
The (ε, δ) pair defining TRC's clipping rule is selected by grid search on CYTO (Table II evaluates average SHD on CYTO), and the same CYTO benchmark is then used to report the 'hitherto best result' SHD of 9/10. The reported number is therefore the best of several configurations tuned on that benchmark, i.e., a fitted performance estimate rather than an independent prediction. The reduction is not complete: the BIC-based reward and ground-truth SHD are external, so the result is not forced by construction, but the headline CYTO figure is partially self-evaluative.
full rationale
The algorithm's objective is the external BIC score (Eq. 4) and all benchmarks are external datasets with ground-truth graphs, so the reported SHD values are not equal by construction to any fitted parameter. The trust-region clipping rule (Eq. 19) is defined directly from the per-edge KL threshold; while the paper's claim that this bounds the joint ratio is mathematically unsupported (a correctness concern, not a circularity), the rule itself is not derived from the result it is used to explain. The only genuine circularity-adjacent issue is that Appendix A selects the (ε, δ) hyperparameters by grid search on the CYTO dataset and Section VI-C then reports TRC's CYTO SHD of 9/10 as a headline result, so this particular number carries a fitting component. This does not make the causal-discovery derivation circular, because the final graph still has to match external ground truth through the BIC reward.
Assumptions & free parameters
free parameters (3)
- epsilon and delta clipping thresholds =
epsilon in {0.1, 0.2}; delta grid searched from 0.05 in steps of 0.015; final chosen pairs not stated globally
- lambda1 and lambda2 penalty schedule =
Table III lists Lambda1 upper bound 0, BICu for lambda2 as -1, Delta1=1, Delta2=10, tu=1000
- Score normalization constant BIC0 =
not reported
assumptions (6)
- domain assumption Additive noise model with independent strictly positive-density noises
- domain assumption Markov and faithfulness assumptions
- domain assumption BIC with linear, quadratic, or GP regression is a valid score
- domain assumption Policy factorization into independent Bernoulli subactions
- ad hoc to paper Per-subaction KL-triggered clipping preserves the joint trust region
- standard math Zhu et al. condition in Equation 20 guarantees best-DAG optimality
Cite this review
Pith. "Pith review of Graph-attention-based Casual Discovery with Trust Region-navigated Clipping Policy Optimization." pith.science (2026). https://pith.science/paper/TPKND6F3
@misc{pith2026241219578,
author = {Pith},
title = {Pith review of: Graph-attention-based Casual Discovery with Trust Region-navigated Clipping Policy Optimization},
year = {2026},
howpublished = {\url{https://pith.science/paper/TPKND6F3}},
note = {Machine review of arXiv:2412.19578}
}
read the original abstract
In many domains of empirical sciences, discovering the causal structure within variables remains an indispensable task. Recently, to tackle with unoriented edges or latent assumptions violation suffered by conventional methods, researchers formulated a reinforcement learning (RL) procedure for causal discovery, and equipped REINFORCE algorithm to search for the best-rewarded directed acyclic graph. The two keys to the overall performance of the procedure are the robustness of RL methods and the efficient encoding of variables. However, on the one hand, REINFORCE is prone to local convergence and unstable performance during training. Neither trust region policy optimization, being computationally-expensive, nor proximal policy optimization (PPO), suffering from aggregate constraint deviation, is decent alternative for combinatory optimization problems with considerable individual subactions. We propose a trust region-navigated clipping policy optimization method for causal discovery that guarantees both better search efficiency and steadiness in policy optimization, in comparison with REINFORCE, PPO and our prioritized sampling-guided REINFORCE implementation. On the other hand, to boost the efficient encoding of variables, we propose a refined graph attention encoder called SDGAT that can grasp more feature information without priori neighbourhood information. With these improvements, the proposed method outperforms former RL method in both synthetic and benchmark datasets in terms of output results and optimization robustness.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
’virus and epidemic’: Causal knowledge activates prediction error circuitry,
D. B. Fenker, M. A. Schoenfeld, M. R. Waldmann, H. Sch ¨utze, H. Heinze, and E. D ¨uzel, “’virus and epidemic’: Causal knowledge activates prediction error circuitry,” J. Cogn. Neurosci., vol. 22, no. 10, pp. 2151–2163, 2010
work page 2010
-
[2]
Causality for machine learning,
B. Sch ¨olkopf, “Causality for machine learning,” CoRR, vol. abs/1911.10500, 2019
arXiv 1911
-
[3]
R. Opgen-Rhein and K. Strimmer, “From correlation to causation networks: a simple approximate learning algorithm and its application to high-dimensional plant gene expression data,” BMC Syst. Biol. , vol. 1, p. 37, 2007
work page 2007
-
[5]
Learning bayesian networks is np-complete,
D. M. Chickering, “Learning bayesian networks is np-complete,” in Learning from Data - Fifth International Workshop on AISTATS 1995. Proceedings, D. Fisher and H. Lenz, Eds. Springer, 1995, pp. 121–130
work page 1995
-
[6]
Causal discovery with reinforcement learning,
S. Zhu, I. Ng, and Z. Chen, “Causal discovery with reinforcement learning,” in 8th International Conference on Learning Representations, ICLR 2020. OpenReview.net, 2020
work page 2020
-
[7]
Estimating the Dimension of a Model,
G. Schwarz, “Estimating the Dimension of a Model,” Annals of Statis- tics, vol. 6, no. 2, pp. 461–464, Jul. 1978
work page 1978
-
[8]
Simple statistical gradient-following algorithms for connectionist reinforcement learning,
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Mach. Learn. , vol. 8, pp. 229– 256, 1992
1992
-
[9]
Policy gradient methods for reinforcement learning with function approxima- tion,
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y . Mansour, “Policy gradient methods for reinforcement learning with function approxima- tion,” in Advances in Neural Information Processing Systems 12 . The MIT Press, 1999, pp. 1057–1063
work page 1999
Show all 48 references
-
[10]
Trust region policy optimization,
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz, “Trust region policy optimization,” in Proceedings of the 32nd International Conference on Machine Learning, ICML 2015 , vol. 37. JMLR.org, 2015, pp. 1889–1897
2015
-
[11]
Proximal policy optimization algorithms,
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” CoRR, vol. abs/1707.06347, 2017
2017 arXiv
-
[12]
Graph attention networks,
P. Velickovic, G. Cucurull, A. Casanova, A. Romero, P. Li `o, and Y . Bengio, “Graph attention networks,” CoRR, vol. abs/1710.10903, 2017
2017 arXiv
-
[13]
Approximating discrete probability distri- butions with dependence trees,
C. K. Chow and C. N. Liu, “Approximating discrete probability distri- butions with dependence trees,” IEEE Trans. Inf. Theory, vol. 14, no. 3, pp. 462–467, 1968
1968
-
[14]
The max-min hill- climbing bayesian network structure learning algorithm,
I. Tsamardinos, L. E. Brown, and C. F. Aliferis, “The max-min hill- climbing bayesian network structure learning algorithm,” Mach. Learn., vol. 65, no. 1, pp. 31–78, 2006
2006
-
[15]
Dags with NO TEARS: continuous optimization for structure learning,
X. Zheng, B. Aragam, P. Ravikumar, and E. P. Xing, “Dags with NO TEARS: continuous optimization for structure learning,” in NeurIPS 2018, 2018, pp. 9492–9503
2018
-
[16]
Generalized score functions for causal discovery,
B. Huang, K. Zhang, and Y . Lin, “Generalized score functions for causal discovery,” in Proceedings of the 24th ACM SIGKDD . ACM, 2018, pp. 1551–1560
2018
-
[17]
Causal discovery with continuous additive noise models,
J. Peters, J. M. Mooij, D. Janzing, and B. Sch ¨olkopf, “Causal discovery with continuous additive noise models,” J. Mach. Learn. Res. , vol. 15, no. 1, pp. 2009–2053, 2014
2009
-
[18]
A machine learning approach to classify pedestrians’ event based on imu and gps,
M. U. Ahmed, S. Brickman, A. Dengg, N. Fasth, M. Mihajlovi ´c, and J. D. Norman, “A machine learning approach to classify pedestrians’ event based on imu and gps,” 2019
2019
-
[19]
Deep learning versus traditional solutions for group trajectory outliers,
A. Belhadi, Y . Djenouri, D. Djenouri, T. Michalak, and J. C. W. Lin, “Deep learning versus traditional solutions for group trajectory outliers,” IEEE Transactions on Cybernetics , pp. 1–12, 2020
2020
-
[20]
Iterative feedback and learning control. servo systems applications,
S. Preitl, R.-E. Precup, P. Zsuzsa, S. Vaivoda, S. Kilyeni, and J. Tar, “Iterative feedback and learning control. servo systems applications,” IFAC Proceedings Volumes (IFAC-PapersOnline), vol. 1, pp. 16–27, 01 2007
2007
-
[21]
Adaptive ekf-based vehicle state estimation with online assessment of local observability,
A. Katriniok and D. Abel, “Adaptive ekf-based vehicle state estimation with online assessment of local observability,” IEEE Trans. Control. Syst. Technol., vol. 24, no. 4, pp. 1368–1381, 2016
2016
-
[22]
Learning functional causal models with generative neural networks,
O. Goudet, D. Kalainathan, P. Caillou, I. Guyon, D. Lopez-Paz, and M. Sebag, “Learning functional causal models with generative neural networks,” Sep. 2017
2017
-
[23]
Structural Agnostic Modeling: Adversarial Learning of Causal Graphs,
D. Kalainathan, O. Goudet, I. Guyon, D. Lopez-Paz, and M. Sebag, “Structural Agnostic Modeling: Adversarial Learning of Causal Graphs,” arXiv e-prints, Mar. 2018
2018
-
[24]
DAG-GNN: DAG structure learning with graph neural networks,
Y . Yu, J. Chen, T. Gao, and M. Yu, “DAG-GNN: DAG structure learning with graph neural networks,” in ICML 2019, 2019
2019
-
[25]
A new model for learning in graph domains,
M. Gori, G. Monfardini, and F. Scarselli, “A new model for learning in graph domains,” in IEEE International Joint Conference on Neural Networks, 2005, vol. 2, 2005, pp. 729–734 vol. 2
2005
-
[26]
The graph neural network model,
F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Monfardini, “The graph neural network model,” IEEE Trans. Neural Networks , vol. 20, no. 1, pp. 61–80, 2009. IEEE TRANSACTIONS ON CYBERNETICS 14
2009
-
[27]
Chebnet: Efficient and stable constructions of deep neural networks with rectified power units using chebyshev approximations,
S. Tang, B. Li, and H. Yu, “Chebnet: Efficient and stable constructions of deep neural networks with rectified power units using chebyshev approximations,” CoRR, vol. abs/1911.05467, 2019
1911 arXiv
-
[28]
Inductive representation learning on large graphs,
W. L. Hamilton, R. Ying, and J. Leskovec, “Inductive representation learning on large graphs,” CoRR, vol. abs/1706.02216, 2017
2017 arXiv
-
[29]
Deep reinforcement learning for multiagent systems: A review of challenges, solutions, and applications,
T. T. Nguyen, N. D. Nguyen, and S. Nahavandi, “Deep reinforcement learning for multiagent systems: A review of challenges, solutions, and applications,” IEEE Transactions on Cybernetics , vol. 50, no. 9, pp. 3826–3839, 2020
2020
-
[30]
Survey of model-based reinforce- ment learning: Applications on robotics,
A. S. Polydoros and L. Nalpantidis, “Survey of model-based reinforce- ment learning: Applications on robotics,” J. Intell. Robotic Syst., vol. 86, no. 2, pp. 153–173, 2017
2017
-
[31]
Online rein- forcement learning control for the personalization of a robotic knee prosthesis,
Y . Wen, J. Si, A. Brandt, X. Gao, and H. H. Huang, “Online rein- forcement learning control for the personalization of a robotic knee prosthesis,” IEEE Transactions on Cybernetics, vol. 50, no. 6, pp. 2346– 2356, 2020
2020
-
[32]
Multitask learning for object localization with deep reinforcement learning,
Y . Wang, L. Zhang, L. Wang, and Z. Wang, “Multitask learning for object localization with deep reinforcement learning,” IEEE Trans. Cogn. Dev. Syst., vol. 11, no. 4, pp. 573–580, 2019
2019
-
[33]
Nonzero-sum game rein- forcement learning for performance optimization in large-scale industrial processes,
J. Li, J. Ding, T. Chai, and F. L. Lewis, “Nonzero-sum game rein- forcement learning for performance optimization in large-scale industrial processes,” IEEE Transactions on Cybernetics, vol. 50, no. 9, pp. 4132– 4145, 2020
2020
-
[34]
Neural Architecture Search with Reinforcement Learning,
B. Zoph and Q. V . Le, “Neural Architecture Search with Reinforcement Learning,” arXiv e-prints,, Nov. 2016
2016
-
[35]
K. A. Bollen, Structural Equations with Latent Variables. Wiley, 1989
1989
-
[36]
Spirtes, C
P. Spirtes, C. Glymour, and R. Scheines, Causation, Prediction, and Search, Second Edition , ser. Adaptive computation and machine learn- ing. MIT Press, 2000
2000
-
[37]
A linear non-gaussian acyclic model for causal discovery,
S. Shimizu, P. O. Hoyer, A. Hyv ¨arinen, and A. J. Kerminen, “A linear non-gaussian acyclic model for causal discovery,” J. Mach. Learn. Res. , vol. 7, pp. 2003–2030, 2006
2003
-
[38]
On the Prop- erties of Neural Machine Translation: Encoder-Decoder Approaches,
K. Cho, B. van Merrienboer, D. Bahdanau, and Y . Bengio, “On the Prop- erties of Neural Machine Translation: Encoder-Decoder Approaches,” arXiv e-prints,, Sep. 2014
2014
-
[39]
Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond,
R. Nallapati, B. Zhou, C. Nogueira dos santos, C. Gulcehre, and B. Xi- ang, “Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond,” arXiv e-prints,, Feb. 2016
2016
-
[40]
Function optimization using connectionist reinforcement learning algorithms,
R. J. Williams and J. Peng, “Function optimization using connectionist reinforcement learning algorithms,” Connection Science, vol. 3, pp. 241– 268, 1991
1991
-
[41]
Prioritized Experience Replay,
T. Schaul, J. Quan, I. Antonoglou, and D. Silver, “Prioritized Experience Replay,” arXiv e-prints,, Nov. 2015
2015
-
[42]
Are deep policy gradient algorithms truly policy gradient algorithms?
A. Ilyas, L. Engstrom, S. Santurkar, D. Tsipras, F. Janoos, L. Rudolph, and A. Madry, “Are deep policy gradient algorithms truly policy gradient algorithms?” CoRR, vol. abs/1811.02553, 2018
2018 arXiv
-
[43]
Ramsey, M
J. Ramsey, M. Glymour, R. Sanchez-Romero, and C. Glymour, “A mil- lion variables and more: the fast greedy equivalence search algorithm for learning high-dimensional graphical causal models, with an application to functional magnetic resonance images,” Int. J. Data Sci. Anal.,...
2017
-
[44]
CAM: causal additive models, high-dimensional order search and penalized regression,
P. B ¨uhlmann, J. Peters, and J. Ernest, “CAM: causal additive models, high-dimensional order search and penalized regression,” CoRR, vol. abs/1310.1533, 2013
2013 arXiv
-
[45]
Gradient- based neural DAG learning,
S. Lachapelle, P. Brouillard, T. Deleu, and S. Lacoste-Julien, “Gradient- based neural DAG learning,” CoRR, vol. abs/1906.02226, 2019
1906 arXiv
-
[46]
Causal Protein-Signaling Networks Derived from Multiparam- eter Single-Cell Data,
K. Sachs, O. Perez, D. Pe’er, D. A. Lauffenburger, and G. P. Nolan, “Causal Protein-Signaling Networks Derived from Multiparam- eter Single-Cell Data,” Science, vol. 308, no. 5721, pp. 523–529, Apr. 2005
2005
-
[47]
Syntren: a generator of synthetic gene expression data for design and analysis of structure learning algorithms,
T. V . den Bulcke, K. V . Leemput, B. Naudts, P. van Remortel, H. Ma, A. Verschoren, B. D. Moor, and K. Marchal, “Syntren: a generator of synthetic gene expression data for design and analysis of structure learning algorithms,” BMC Bioinform., vol. 7, p. 43, 2006
2006
-
[48]
Truly proximal policy optimization,
Y . Wang, H. He, and X. Tan, “Truly proximal policy optimization,” in UAI 2019 , ser. Proceedings of Machine Learning Research, vol. 115. AUAI Press, pp. 113–122
2019
-
[49]
A hybrid method for nonlinear equations,
M. J. D. Powell, “A hybrid method for nonlinear equations,” in Nu- merical Methods for Nonlinear Algebraic Equations , P. Rabinowitz, Ed. Gordon and Breach, 1970. Shixuan Liu received the B.S. degree in Systems Engineering in 2019 from National University of Defense Technology...
1970
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.