REVIEW 2 major objections 4 minor 65 references
Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence
T0 review · 2 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Two instances of one model co-fail on 90.0% of missions where either fails, so the independence assumption behind compositional reliability bounds gives way, and the paper responds with a certificate that assumes no dependence structure.
desk verdict A serious, unusually transparent empirical and theoretical package on correlated failures in multi-agent systems; the 90% co-failure finding is solid modulo one legitimate scorer-validity concern, and the moment-set certificate is a real contribution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two objects carry the argument. Diagnostically, the signed compositional gap: for two components, the true joint failure probability exceeds the independence product by exactly $Cov(h_1, h_2)$, the covariance of the two hard verdicts, so positive dependence always over-credits redundancy and the direction of the error is fixed without further estimation. Constructively, the moment-set certificate: a linear program over the $2^m$ cells of the joint law on $\{0,1\}^m$, minimizing the all-success probability subject to the constraint that measured co-execution moments — single-stage successes, pairwise co-successes, and optionally triple co-successes — lie inside a Bonferroni–Clopper–Pearson box around their empirical values; the box covers by the union bound, the true law is feasible whenever it covers, and the LP minimum is the certified floor. A companion betting e-process, $E_R = \prod_{r \le R}\bigl(1 + \lambda_r (y_r - p_0)\bigr)$, gives an anytime-valid certificate whose null constrains only a conditional mean, so it needs no independence assumption at all.
What would settle it
Rescore a random sample of the confirmatory missions with a contract set and scoring code written independently of the paper's authors, and compare the same-model co-failure rate of 90.0% and the log-odds contrasts; if the co-failure pattern moves materially the headline is a scoring artifact, while reproduction would confirm genuine model dependence. A sharper variant runs the substitution with two models matched on marginal failure rate but from different lineages, separating shared competence from shared inductive bias.
Extended reading notes
Core claim
The paper's claim is that the conditional-independence condition C5 of compositional contract theory fails for same-model composition, that the failure is large and signed, and that it can be certified around without any dependence assumption. In the preregistered confirmatory arm, two instances of mistral-small-24b in a two-agent handoff co-fail on 2177 of the 2418 missions on which either fails — 90.0%, against the 14.6% the independence product predicts — with log odds ratio 6.66 (95% CI [6.38, 7.00]) and $\phi = 0.916$. Substituting a different model into the second agent reduces the association significantly in six of six contrasts across three topologies; substituting a different vendor, with the model already different, does not, a registered hypothesis reported as a null. The paper further shows that the assumption-free Fréchet–Hoeffding sandwich, which brackets a joint probability from marginals alone, is vacuous — its certified floor is zero whenever mean component reliability falls below $1 - 1/m$ — and that a certificate built on a fitted dependence model loses coverage of the truth as the sample grows, because the identification gap is $O(1)$ while the bootstrap haircut is $O(n^{-1/2})$. Its remedy is a linear program over the joint law on $\{0,1\}^m$, minimized over all distributions whose co-execution moments lie in a Bonferroni–Clopper–Pearson box: valid with no dependence assumption, sharp for the moments supplied, and monotone in the moment family, and on four-stage data enriching ten moment functionals to fourteen narrows the identified interval by 85.7% and lifts the certified floor from 0.2455 to 0.4116.
Load-bearing premise
The headline result stands on the deterministic gold scoring code measuring contract compliance correctly, even though the same authors wrote both the contracts and the scorer; a systematic scoring bias would make the co-failure counts and arm contrasts artifacts rather than true model dependence.
Editorial extensions
If this is right
- At the measured dependence the independence product over-credits redundancy by a wide margin: the two agents would co-fail on 14.6% of missions if independent and actually co-fail on 36.3%, so a dashboard that multiplies reliabilities reports a reassuring number whose governing assumption the data reject.
- Model identity, not vendor identity, is the axis on which to choose redundancy: substituting a different model reduced the association significantly in six of six contrasts across three topologies, while substituting a different vendor, with the model already different, produced no consistent reduction.
- A certificate built on a fitted dependence model is worse than no certificate: its bootstrap lower bound loses coverage of the true reliability as the sample grows, so more data narrows an interval around the wrong target with no visible symptom.
- Operators can tighten the certified floor without any dependence assumption by measuring more co-execution moments: on four-stage data, enriching ten moment functionals to fourteen narrowed the identified interval by 85.7% and lifted the floor from 0.2455 to 0.4116, and under pre-allocated Bonferroni spending the tightening is monotone by construction.
- Marginal-bounded dependence statistics — Jaccard, $\phi$, Kendall's $\tau_a$ — can reverse the apparent ordering of conditions when the compared agents fail at different rates, so a marginal-free statistic such as the log odds ratio should be reported alongside them.
Reading between the lines
- If the co-failure rate transfers beyond the retail and financial task domains tested, any safety case built on same-model redundant review is effectively relying on a single point of failure: a reviewer and a writer drawn from one model supply far less independent evidence than the design assumes, and the honest redundancy count may be one rather than two.
- The coverage-collapse theorem is not specific to agents: any bootstrap interval built on a misspecified parametric model with a fixed identification gap will, past some sample size, sit entirely off the truth, which cautions against model-based certificates in copula-based risk and reliability practice generally.
- Because the anytime-valid e-process needs no independence assumption, it is the one certificate in the paper that survives the failure the paper documents; a natural deployment pattern is a continuously re-earned certificate that a team may watch and stop on at bounded type-I cost.
- A testable extension suggested by the paper's own limitation section: run the substitution design with two models matched on marginal failure rate but from different lineages, which would separate the contribution of shared competence from shared inductive bias to the co-failure signal.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper tests the conditional-independence condition C5 that licenses multiplicative compositional reliability bounds for multi-agent systems. In a preregistered campaign of 18,000 two-agent handoff missions scored by deterministic code, two instances of mistral-small-24b co-fail on 90.0% of missions on which either fails (log OR = 6.66, 95% CI [6.38, 7.00]); substituting a different model significantly reduces the association in six of six contrasts, while substituting a different vendor, with the model already different, does not. The paper then develops a finite-sample reliability certificate based on a linear program over measured co-execution moments with a Bonferroni-Clopper-Pearson box, proves it sound and sharp, adds an anytime-valid e-process certificate, and proves that bootstrap bounds on a fitted dependence model lose coverage of the true reliability as the sample grows. The evaluation also reports a negative control, cross-backend replication, and an ablation of the i.i.d. assumption. The paper is unusually transparent about its own limitations, including the post-hoc choice of the log odds ratio for H2 and the repository-based rather than external preregistration timestamp.
Significance. If the empirical finding holds, it is significant: it provides controlled evidence that the conditional-independence assumption fails for same-model composition, with a signed, practically important consequence for redundant agent designs. The theoretical machinery is a genuine contribution: the moment-set LP certificate of Theorem 5.2 is a sound finite-sample, copula-agnostic bound, sharp for the supplied moments, and the coverage-collapse result of Theorem 4.2 is a useful warning about fitted-dependence certificates. The paper ships complete proofs, released code, deterministic scoring, and regenerable statistics, which are strengths. The anytime-valid certificate is standard machinery but is applied cleanly and with careful empirical checking. The main risks are empirical rather than theoretical: the headline co-failure numbers depend entirely on the authors' own deterministic scorer, and the substitution manipulation confounds model identity with model capability; both are acknowledged but not resolved.
major comments (2)
- [Sections 10.1, 10.2.1, 11.3] The headline co-failure estimate (J = 0.9003, log OR = 6.66) and the arm contrasts of Table 3 are produced entirely by the authors' deterministic gold scorer. Because mission identifiers embed the condition name and each arm draws its own missions, a scorer rule that misfires on particular mission or output types can inflate co-failure counts in one arm more than another, so the Section 11.3 defense that the same contracts and scorer are used in all arms does not fully secure the differential claim. The internal negative control of Table 4 holds the compared pair fixed and therefore cannot validate the scorer's behavior across different model identities. An independent scorer or human rescoring of a stratified subsample is needed before the 90% co-failure finding is relied on.
- [Sections 10.1, 10.2.2, 11.3] The manipulation confounds model identity with model capability: the same_vendor substitution uses ministral-8b, a weaker model, and the different_vendor arm uses gemma-3-12b-it. The six significant same_model-versus-substitution contrasts, and the Section 10.9 conclusion that the operative variable is the model not the vendor, could in part be driven by capability-correlated failure modes rather than by shared model weights. The marginal-free statistics remove the effect of different marginal failure rates but not the effect of capability-correlated failure modes, and the negative control does not address this because the control pair is never substituted. A design with models matched on failure rate, or a direct manipulation of shared weights, is required to support the causal attribution in the abstract.
minor comments (4)
- [Equation (25), Definition 3.17] The quantity 2(p11p00 − p10p01) is labelled Kendall's τa, but the conventional sample tau-a for binary variables is a factor of 2 larger (approximately 4(p11p00 − p10p01) with a finite-population correction). Please state the scaling convention explicitly, since the registered H1 threshold is stated on τa.
- [Abstract and Section 10.2.2] The phrase 'a registered hypothesis reported as a null' for the same_vendor versus different_vendor comparison should be qualified: the registered H2 was a three-level ordering on τa, and the vendor-level comparison is a post-hoc decomposition of that ordering. The paper's Section 11.3 disclosure is accurate, but the abstract's wording invites the reading that the vendor null was itself preregistered.
- [Section C.4] The preregistration timestamp rests on repository history rather than on an external registry. The paper discloses this, but given the confirmatory role of the registration, an independent timestamp or third-party registration would materially strengthen the claim.
- [Table 3 caption] The verdict label 'conflict' in the SV−DV rows is not defined in the caption; please state that it means the marginal-sensitive and marginal-free statistics disagree in sign or significance, as explained in the text.
Circularity Check
No circular derivation: the empirical co-failure result is a direct preregistered measurement and the certificate theorems are proved from stated moment constraints.
full rationale
I walked the derivation chain from the measured co-failure tables (Table 2) to the headline 90.0% overlap, the arm contrasts (Table 3), and the moment-set and anytime-valid certificates (Theorems 5.2, 6.1, 7.1). The co-failure percentage is arithmetic on n11/(n11+n10+n01) from logs scored by deterministic code; it is not a fitted parameter renamed as a prediction. The LP certificate minimizes all-success over a Clopper-Pearson moment box and is proved sound; sharpness and monotonicity are proven, and the Bonferroni allocation that makes monotonicity hold by construction is explicitly disclosed (Proposition 6.2). Theorem 4.2's coverage-collapse result is a delta>0 misspecification argument, and E3's witness is explicitly constructed as the LP minimizer that is moment-indistinguishable from the Gaussian law; this is a valid illustration of the theorem, not an independent empirical prediction, and the paper says so. The only self-citation is to the v1 paper for the ABC framework definitions and carried-forward single-agent evidence (Sections 2.1, 10.8); the framework formulas are restated in full in Section 3, and the v1 empirical evidence is not used to derive any new result, so it is not load-bearing circularity. The paper itself flags the main validity threats: the gold scorer shares an author with the contracts (Section 11.3), the H2 estimator was chosen after seeing the reversal (Section 11.3), the preregistration timestamp rests on repository history rather than an external registry (Section C.4), and the Meta breadth arm suffered outcome-dependent attrition (Section 10.6). These are correctness and validity concerns, not circularity: none of them makes a stated prediction equal to an input by construction.
Assumptions & free parameters
assumptions (4)
- domain assumption Missions in each arm are i.i.d. draws from the generator distribution.
- domain assumption The deterministic gold scoring code correctly evaluates contract compliance.
- domain assumption The v1 ABC framework (Definitions 3.1-3.13) is taken as given, including the (p, delta, k)-satisfaction notion and the drift dynamics.
- domain assumption Composition rules (series, quorum, parallel) are deterministic; aggregator nodes are code, not models.
Cite this review
Pith. "Pith review of Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence." pith.science (2026). https://pith.science/paper/JWJI2PM6
@misc{pith2026260812895,
author = {Pith},
title = {Pith review of: Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence},
year = {2026},
howpublished = {\url{https://pith.science/paper/JWJI2PM6}},
note = {Machine review of arXiv:2608.12895}
}
read the original abstract
Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that is routinely stated and rarely tested. We test it. Two instances of one model, in a two-agent handoff, co-fail on 90.0% of the missions on which either fails (log OR 6.66, 95% CI [6.38, 7.00]; phi 0.916), in a preregistered evaluation of 18,000 missions scored by deterministic code with no LLM judge. Substituting a different model reduces the association in six of six contrasts; substituting a different vendor, model already different, does not -- a registered hypothesis reported as a null. The error is signed and runs against the operator: positive dependence inflates joint failure above the independence product, so redundancy is over-credited exactly when components share a model. The assumption-free alternative is often vacuous, and fitting a dependence model is worse: we prove a bootstrap bound on a fitted model's functional loses coverage of the truth as n grows, the identification gap being O(1) while the bootstrap haircut is O(n^{-1/2}). More data makes such a certificate worse, with no visible symptom. We give a finite-sample certificate assuming no dependence structure: a linear program over the joint, over a Bonferroni-Clopper-Pearson box around measured co-execution moments. It is sound, sharp for the information supplied, and monotone in the moment family. Enriching ten moment functionals to fourteen narrows the identified interval by 85.7% and lifts the certified floor from 0.2455 to 0.4116. A companion anytime-valid certificate holds type-I error at 0.0471 under optional stopping. Common dependence statistics are marginal-bounded and can reverse an apparent ordering of conditions when the compared agents fail at different rates. Contracts, scoring code, analysis scripts, and the preregistration are released.
Reference graph
Works this paper leans on
-
[1]
Constitutional AI : Harmlessness from AI feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, et al. Constitutional AI : Harmlessness from AI feedback. arXiv preprint arXiv:2212.08073, 2022
arXiv 2022
-
[2]
Richard E. Barlow and Frank Proschan. Statistical Theory of Reliability and Life Testing. Holt, Rinehart and Winston, 1975
work page 1975
- [3]
-
[4]
International AI safety report 2026
Yoshua Bengio et al. International AI safety report 2026. arXiv preprint arXiv:2602.21012, 2026
arXiv 2026
-
[5]
Can AI agents agree? arXiv preprint arXiv:2603.01213, 2026
Fr\'ed\'eric Berdoz, Leonardo Rugli, and Roger Wattenhofer. Can AI agents agree? arXiv preprint arXiv:2603.01213, 2026
arXiv 2026
-
[6]
Optimal inequalities in probability theory: A convex optimization approach
Dimitris Bertsimas and Ioana Popescu. Optimal inequalities in probability theory: A convex optimization approach. SIAM Journal on Optimization, 15 0 (3): 0 780--804, 2005
work page 2005
-
[7]
Varun Pratap Bhardwaj. Agent behavioral contracts: Formal specification and runtime enforcement for reliable autonomous AI agents. arXiv preprint arXiv:2602.22302, 2026
arXiv 2026
-
[8]
Carlo E. Bonferroni. Teoria statistica delle classi e calcolo delle probabilit\`a. Pubblicazioni del R. Istituto Superiore di Scienze Economiche e Commerciali di Firenze, 8: 0 3--62, 1936
work page 1936
Show all 65 references
-
[9]
An Investigation of the Laws of Thought
George Boole. An Investigation of the Laws of Thought. Walton and Maberly, 1854
-
[10]
LangChain
Harrison Chase. LangChain . https://github.com/langchain-ai/langchain, 2022
2022
-
[11]
Clarke, Orna Grumberg, and Doron A
Edmund M. Clarke, Orna Grumberg, and Doron A. Peled. Model Checking. MIT Press, 1999
1999
-
[12]
C. J. Clopper and E. S. Pearson. The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika, 26 0 (4): 0 404--413, 1934
1934
-
[13]
Abstract interpretation: A unified lattice model for static analysis of programs by construction or approximation of fixpoints
Patrick Cousot and Radhia Cousot. Abstract interpretation: A unified lattice model for static analysis of programs by construction or approximation of fixpoints. In POPL, pages 238--252, 1977
1977
-
[14]
Tenenbaum, and Igor Mordatch
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. Improving factuality and reasoning in language models through multiagent debate. In ICML, 2024
2024
-
[15]
Bootstrap methods: Another look at the jackknife
Bradley Efron. Bootstrap methods: Another look at the jackknife. The Annals of Statistics, 7 0 (1): 0 1--26, 1979
1979
-
[16]
Endres and Johannes E
Dominik M. Endres and Johannes E. Schindelin. A new metric for probability distributions. IEEE Transactions on Information Theory, 49 0 (7): 0 1858--1860, 2003
2003
-
[17]
Ernst, Jeff H
Michael D. Ernst, Jeff H. Perkins, Philip J. Guo, Stephen McCamant, Carlos Pacheco, Matthew S. Tschantz, and Chen Xiao. The Daikon system for dynamic detection of likely invariants. Science of Computer Programming, 69 0 (1--3): 0 35--45, 2007
2007
-
[18]
Sur les tableaux de corr\'elation dont les marges sont donn\'ees
Maurice Fr\'echet. Sur les tableaux de corr\'elation dont les marges sont donn\'ees. Annales de l'Universit\'e de Lyon, Section A, 14: 0 53--77, 1951
1951
-
[19]
The statistical crisis in science
Andrew Gelman and Eric Loken. The statistical crisis in science. American Scientist, 102 0 (6): 0 460--465, 2014
2014
-
[20]
Safe testing
Peter Gr\"unwald, Rianne de Heide, and Wouter Koolen. Safe testing. Journal of the Royal Statistical Society B, 86 0 (5): 0 1091--1128, 2024
2024
-
[21]
Best possible inequalities for the probability of a logical function of events
Theodore Hailperin. Best possible inequalities for the probability of a logical function of events. The American Mathematical Monthly, 72 0 (4): 0 343--359, 1965
1965
-
[22]
Multi-agent risks from advanced AI
Lewis Hammond et al. Multi-agent risks from advanced AI . arXiv preprint arXiv:2502.14143, 2025
2025 arXiv
-
[23]
C. A. R. Hoare. An axiomatic basis for computer programming. Communications of the ACM, 12 0 (10): 0 576--580, 1969
1969
-
[24]
ur Angewandte Mathematik der Universit\
Wassily Hoeffding. Ma stabinvariante K orrelationstheorie. Schriften des Mathematischen Instituts und des Instituts f\"ur Angewandte Mathematik der Universit\"at Berlin, 5: 0 179--233, 1940
1940
-
[25]
Probability inequalities for sums of bounded random variables
Wassily Hoeffding. Probability inequalities for sums of bounded random variables. Journal of the American Statistical Association, 58 0 (301): 0 13--30, 1963
1963
-
[26]
MetaGPT : Meta programming for a multi-agent collaborative framework
Sirui Hong, Mingchen Zhuge, Jonathan Chen, et al. MetaGPT : Meta programming for a multi-agent collaborative framework. In ICLR, 2024
2024
-
[27]
Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon
Steven R. Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon. Time-uniform, nonparametric, nonasymptotic confidence sequences. The Annals of Statistics, 49 0 (2): 0 1055--1080, 2021
2021
-
[28]
Counterfactual graph for multi-agent LLM calibration
Jiatan Huang, Mingchen Li, Ziming Li, Sunjae Kwon, Hong Yu, and Chuxu Zhang. Counterfactual graph for multi-agent LLM calibration. arXiv preprint arXiv:2605.30653, 2026
2026 arXiv
-
[29]
The distribution of the flora in the alpine zone
Paul Jaccard. The distribution of the flora in the alpine zone. New Phytologist, 11 0 (2): 0 37--50, 1912
1912
-
[30]
Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. SWE-bench : Can language models resolve real-world GitHub issues? In ICLR, 2024
2024
-
[31]
Maurice G. Kendall. A new measure of rank correlation. Biometrika, 30 0 (1--2): 0 81--93, 1938
1938
-
[32]
Survey Sampling
Leslie Kish. Survey Sampling. John Wiley and Sons, New York, 1965
1965
-
[33]
Specifying Systems: The TLA+ Language and Tools for Hardware and Software Engineers
Leslie Lamport. Specifying Systems: The TLA+ Language and Tools for Hardware and Software Engineers . Addison-Wesley, 2002
2002
-
[34]
Rustan M
K. Rustan M. Leino. Dafny: An automatic program verifier for functional correctness. In LPAR, pages 348--370, 2010
2010
-
[35]
Encouraging divergent thinking in large language models through multi-agent debate
Tian Liang, Zhiwei He, Wenxiang Jiao, et al. Encouraging divergent thinking in large language models through multi-agent debate. In EMNLP, 2024
2024
-
[36]
Divergence measures based on the S hannon entropy
Jianhua Lin. Divergence measures based on the S hannon entropy. IEEE Transactions on Information Theory, 37 0 (1): 0 145--151, 1991
1991
-
[37]
AgentBench : Evaluating LLMs as agents
Xiao Liu, Hao Yu, Hanchen Zhang, et al. AgentBench : Evaluating LLMs as agents. In ICLR, 2024
2024
-
[38]
Shay Seiya McDonnell, Avantika Singh, Quoc-Viet Pham, Vratislav Havlik, and Gregory M. P. O'Hare. Harnessing disagreement: Detecting correlated agreement blindness in multi-agent triage. arXiv preprint arXiv:2607.19899, 2026. Accepted, PAAMS 2026
2026 arXiv
-
[39]
Applying ``design by contract''
Bertrand Meyer. Applying ``design by contract''. Computer, 25 0 (10): 0 40--51, 1992
1992
-
[40]
Roger B. Nelsen. An Introduction to Copulas. Springer, 2nd edition, 2006
2006
-
[41]
Nosek, Charles R
Brian A. Nosek, Charles R. Ebersole, Alexander C. DeHaven, and David T. Mellor. The preregistration revolution. Proceedings of the National Academy of Sciences, 115 0 (11): 0 2600--2606, 2018
2018
-
[42]
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, et al. Training language models to follow instructions with human feedback. In NeurIPS, 2022
2022
-
[43]
O'Brien, Carrie J
Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In UIST, 2023
2023
-
[44]
ChatDev : Communicative agents for software development
Chen Qian, Wei Liu, Hongzhang Liu, et al. ChatDev : Communicative agents for software development. In ACL, 2024
2024
-
[45]
VerifyMAS : Hypothesis verification for failure attribution in LLM multi-agent systems
Hezhe Qiao, Hanghang Tong, Ee-Peng Lim, Bing Liu, and Guansong Pang. VerifyMAS : Hypothesis verification for failure attribution in LLM multi-agent systems. arXiv preprint arXiv:2605.17467, 2026
2026 arXiv
-
[46]
Governed capability evolution: Lifecycle-time compatibility checking and rollback for AI -component-based systems
Xue Qin, Simin Luan, John See, Zeyd Boukhers, Cong Yang, and Zhijun Li. Governed capability evolution: Lifecycle-time compatibility checking and rollback for AI -component-based systems. arXiv preprint arXiv:2604.08059, 2026
2026 arXiv
-
[47]
FALAT : Tracing failures in LLM agent trajectories via dependency-guided search
Md Nakhla Rafi, Md Ahasanuzzaman, Dong Jae Kim, Zhijie Wang, and Tse-Hsun Chen. FALAT : Tracing failures in LLM agent trajectories via dependency-guided search. arXiv preprint arXiv:2606.00765, 2026
2026 arXiv
-
[48]
Game-theoretic statistics and safe anytime-valid inference
Aaditya Ramdas, Peter Grünwald, Vladimir Vovk, and Glenn Shafer. Game-theoretic statistics and safe anytime-valid inference. Statistical Science, 38 0 (4): 0 576--601, 2023
2023
-
[49]
NeMo guardrails: A toolkit for controllable and safe LLM applications with programmable rails
Traian Rebedea, Razvan Dinu, Makesh Sreedhar, Christopher Parisien, and Jonathan Cohen. NeMo guardrails: A toolkit for controllable and safe LLM applications with programmable rails. In EMNLP System Demonstrations, 2023
2023
-
[50]
Statistical methods related to the law of the iterated logarithm
Herbert Robbins. Statistical methods related to the law of the iterated logarithm. The Annals of Mathematical Statistics, 41 0 (5): 0 1397--1409, 1970
1970
-
[51]
Testing by betting: A strategy for statistical and scientific communication
Glenn Shafer. Testing by betting: A strategy for statistical and scientific communication. Journal of the Royal Statistical Society A, 184 0 (2): 0 407--431, 2021
2021
-
[52]
Game-Theoretic Foundations for Probability and Finance
Glenn Shafer and Vladimir Vovk. Game-Theoretic Foundations for Probability and Finance. Wiley, 2019
2019
-
[53]
Reflexion : Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion : Language agents with verbal reinforcement learning. In NeurIPS, 2023
2023
-
[54]
Simmons, Leif D
Joseph P. Simmons, Leif D. Nelson, and Uri Simonsohn. False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22 0 (11): 0 1359--1366, 2011
2011
-
[55]
Fonctions de r\'epartition \`a n dimensions et leurs marges
Abe Sklar. Fonctions de r\'epartition \`a n dimensions et leurs marges. Publications de l'Institut de Statistique de l'Universit\'e de Paris, 8: 0 229--231, 1959
1959
-
[56]
The one-sided barrier problem for G aussian noise
David Slepian. The one-sided barrier problem for G aussian noise. Bell System Technical Journal, 41 0 (2): 0 463--501, 1962
1962
-
[57]
\'Etude critique de la notion de collectif
Jean Ville. \'Etude critique de la notion de collectif. Gauthier-Villars, Paris, 1939
1939
-
[58]
Sequential tests of statistical hypotheses
Abraham Wald. Sequential tests of statistical hypotheses. The Annals of Mathematical Statistics, 16 0 (2): 0 117--186, 1945
1945
-
[59]
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, et al. Self-consistency improves chain of thought reasoning in language models. In ICLR, 2023
2023
-
[60]
Estimating means of bounded random variables by betting
Ian Waudby-Smith and Aaditya Ramdas. Estimating means of bounded random variables by betting. Journal of the Royal Statistical Society B, 86 0 (1): 0 1--27, 2024
2024
-
[61]
AutoGen : Enabling next-gen LLM applications via multi-agent conversation
Qingyun Wu, Gagan Bansal, Jieyu Zhang, et al. AutoGen : Enabling next-gen LLM applications via multi-agent conversation. In COLM, 2024
2024
-
[62]
ReAct : Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct : Synergizing reasoning and acting in language models. In ICLR, 2023
2023
-
[63]
Udny Yule
G. Udny Yule. On the association of attributes in statistics: With illustrations from the material of the childhood society. Philosophical Transactions of the Royal Society A, 194: 0 257--319, 1900
1900
-
[64]
Rethinking the reliability of multi-agent system: A perspective from Byzantine fault tolerance
Lifan Zheng, Jiawei Chen, Qinghong Yin, Jingyuan Zhang, Xinyi Zeng, and Yu Tian. Rethinking the reliability of multi-agent system: A perspective from Byzantine fault tolerance. arXiv preprint arXiv:2511.10400, 2025
2025
-
[65]
Xu, Hao Zhu, et al
Shuyan Zhou, Frank F. Xu, Hao Zhu, et al. WebArena : A realistic web environment for building autonomous agents. In ICLR, 2024
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.