{"id":"fd0e3ac1-c0bf-4ff3-bfef-fbd7676d0ac1","arxiv_id":"2608.10030","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"AEROBAT, an LLM-based multi-agent system, automates the full pipeline of behavioral research on AI agents and reports moderate-to-strong evidence for 26 of 79 tested hypotheses.","lead":"A new multi-agent system called AEROBAT automatically runs behavioral experiments on AI agents, from hypothesis generation to written reports. It tested 79 hypotheses across 12 behaviors and found statistical support for 26 of them.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central empirical claim depends on LLM-generated behavior scores that are never validated against independent human ratings; the 26 findings may be artifacts of the rubric rather than properties of the target behaviors.","rationale":"I agree with the reader's diagnosis. The score-validity gap is the single most load-bearing concern because it is prior to every empirical result: without criterion validity for y_hat, the 26 findings are not findings about behavior. The paper does provide real independent support elsewhere—environment fidelity checks (Section 3.3), cross-model generalization (Section 3.2), and external consistency with prior literature (Appendix F)—but none of these substitutes for human labels. Internal consistency is necessary, not sufficient; proxy consistency with prior work is indirect and mostly confined to well-known effects; novel findings remain unvalidated. The concrete test above would settle the issue. If it passes, the verdict could be strengthened; if it fails, the central empirical claim collapses. Multiple testing and manual hypothesis selection are secondary: they affect the strength of the evidence but are addressable, whereas an invalid dependent variable would undermine every downstream conclusion. Because this concern is exactly what the reader flagged, and the appropriate remedy is external validation rather than rejection, CONDITIONAL remains the right verdict.","tokens_in":54448,"tokens_out":3409,"duration_ms":37561,"concrete_test":"Sample 40 simulation transcripts stratified across the 12 behaviors, covering significant, null, and inconclusive hypotheses. Have two independent human annotators rate each transcript for the target behavior on a global scale using only the behavior definition ydef (not the LLM rubric), and compute agreement (ICC, weighted kappa) with AEROBAT's aggregate y_hat. Then re-run the Bayesian monotone-increment model (Eq. 8) for the hypotheses represented in the sample using human scores as the dependent variable. If the human–LLM correlation is below about 0.5, or if the sign or class of the effect changes for any headline finding, the 26 findings are not established as findings about the target behaviors.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim—that AEROBAT found moderate-to-strong evidence for 26 of 79 hypotheses about real target behaviors such as deception, empathy, and sycophancy—is only as strong as the Stage 4 dependent variable. Every Bayes factor and effect size in Fig. 3 is computed from y_hat in Eq. (7), the blind reviewer's rubric-based score, with no validation against human judgments of the same transcripts. Appendix E.1 is an internal-consistency analysis: it shows the five evidence-class scores cohere (mean alpha 0.90) and are semantically distinct, but coherence does not establish criterion validity. Section 3.3 validates that the environments instantiate the intended manipulations, not that the reviewer's scores track the target construct. Appendix F compares findings to prior literature, but it is a post-hoc, author-coded mapping of 25 resolved hypotheses, mostly 'proxy consistency,' and cannot validate the novel findings that make the system valuable. Consequently, if the LLM reviewer encodes text-surface patterns—e.g., words like 'strategic,' 'selective,' or conflict framing—rather than the construct, then the monotone-increment model is detecting a reliable but construct-irrelevant signal. The causal manipulations could shift y_hat without shifting the target behavior, or vice versa. This is the load-bearing assumption: the headline claim is about behavior, not about the reviewer's score.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"AEROBAT is a multi-agent LLM system that, given a user-specified target behavior Y, claims to automate the entire behavioral-science pipeline for AI agents: Stage 1 generates a behavioral definition, a multi-class scoring rubric, hypotheses about causal variables X, and domains; Stage 2 designs matched environment configurations varying only X, with multiple textual realizations per level; Stage 3 runs multi-round simulations via a simulator agent and a subject agent; Stage 4 has a blind reviewer score the subject agent's behavior against the rubric; and a research-manager agent gates each stage and writes a final report. Statistical analysis uses a Bayesian monotone-increment model (Eq. 8) with a closed-form Bayes factor and standardized effect size, plus a block-stratified Kendall's tau check. In experiments with 12 target behaviors and GPT-5-mini as the subject agent, the system generated 79 hypotheses, ran 1,240 experiments and 23,512 simulation rounds, and reported 26 hypotheses with BF10 >= 3, including two extended example reports (instruction divergence -> literal instruction-following; goal conflict -> deception). Additional analyses cover cross-subject-agent generalization (Sec. 3.2), environment fidelity (Sec. 3.3), rubric internal consistency and robustness (App. E.1), prior sensitivity and Monte Carlo error (App. E.3), gating statistics and cost (Apps. E.2, E.4), and a comparison of 25 resolved findings to prior literature (App. F).","tokens_in":2469,"tokens_out":2683,"duration_ms":240228,"significance":"If the findings hold up, this is a meaningful step for AI-agent behavioral science: it is the first system I am aware of that executes the full controlled-experiment cycle, including hypotheses, matched designs, multiple realizability, blind assessment, analysis, and writing, for arbitrary target behaviors, and it does so at scale. The paper's strengths should be credited: the environment model (control, parametrization, multiple realizability) is well designed; the statistical layer is carefully specified with a closed-form Bayes factor tailored to the monotone hypothesis space, prior-scale and Monte-Carlo-error sensitivity analyses, and a model-free rank check; the environment-fidelity evaluation (three inverse problems plus human ratings) is more thorough than typical for LLM-agent papers; and the rubric diagnostics (internal consistency, semantic specificity, null-score and single-class robustness) are a serious attempt at measurement quality. However, the validity of the central empirical claim is conditional on an unvalidated link: the Stage-4 behavior score y-hat.","major_comments":[{"comment":"The 26 headline findings in Fig. 3 are statements about the Stage-4 review score y-hat, the mean of evidence-class scores produced by a GPT-5.1 reviewer against a rubric that GPT-5.1 generated in Stage 1; the paper never validates y-hat against independent human ratings of the same simulation transcripts. This is load-bearing because the Abstract and Section 3.1 present the findings as being about behavioral constructs, such as deception, empathy, and sycophancy, rather than about a model-generated rating. Appendix E.1 establishes internal consistency (mean alpha = 0.90) and semantic specificity of the rubric text, and Section 3.3 shows that environments instantiate the intended manipulations; neither establishes that y-hat tracks the target construct rather than surface text features. The risk is concrete in the worked example: the extreme-condition role prompt (App. A.3) uses wording such as 'selective,' 'narrower framing,' and 'actively look for reasonable frames,' and the reviewer's 'strategic intent cues' score of 1 (App. A.5) is justified by exactly this kind of message-shaping language; since the same model family produces the rubric, the manipulation text, and the review, vocabulary overlap is plausible. I ask for a criterion-validity study: human annotators, ideally plus a reviewer from a different model family, should score a stratified sample of transcripts spanning behaviors and causal-variable levels against the Stage-1 rubrics, with agreement statistics (ICC or weighted kappa) reported, and preferably an analysis showing that the LLM scores predict human labels after controlling for manipulation-wording cues. If the authors intend y-hat to be the object of study (behavior-as-judged-by-the-LLM), the paper should be reframed accordingly; as written, it claims findings about behavior.","section":"Sec. 3.1; Eq. (7); Apps. E.1, A.3, A.5"},{"comment":"The Stage-4 reviewer is blind to the hypothesis and to other conditions, but not to the within-run manipulation. Its input is fij.init (the subject agent's system prompt) plus the full run history; for manipulations mapped to f.roles, f.authority, or f.constraints, the reviewer reads the manipulated text directly in the agent's system prompt, and for world/consequence components it reads the manipulation in the rendered passage. Because the configuration designer receives the rubric (Eqs. (2)-(4)) and both roles are played by GPT-5.1, scoring criteria can become textually aligned with the manipulation language: the extreme-conflict role prompt in App. A.3 instructs the agent to be 'selective' and to 'actively look for reasonable frames,' and the reviewer's strategic-intent score in App. A.5 is justified by precisely this kind of wording. The instruction that the target of evaluation is the subject agent, not the environment or actors, mitigates but does not test this anchoring path. The issue is not peripheral: prompt-embedded components account for 16 of the 26 significant hypotheses (Fig. 4-left: f.roles 11/13, f.authority 3/13, f.constraints 2/4). I request an analysis that separates anchoring from behavior-based scoring, for example by scoring a sample of transcripts with and without the system prompt (or with a neutralized prompt) and by comparing the LLM reviewer's scores with human scores based on the agent's actions alone.","section":"Sec. 2.2 Stage 4 (Eq. (7)); App. A.5"},{"comment":"Appendix F is offered as evidence of overall validity ('These results indirectly show the overall validity of AEROBAT's research pipeline'), but as written it cannot carry that weight. Of the 29 resolved hypotheses, only 5 have near-direct matches to prior work, and one of those is a direct inconsistency (resource scarcity level -> compete, finding no effect where ALYMPICS [33] found increased competitive bidding under scarcity). The other 19 'proxy consistency' classifications use a loose matching criterion ('conceptually similar behavior-cause pair, with materially different manipulation, domain, configuration, or behavioral measure') with no pre-specified protocol and no inter-rater reliability. Consequently, the abstract's 'including some novel ones' refers precisely to the subset of findings with no external anchor, whose validity depends entirely on the unvalidated Stage-4 score. The authors should either strengthen this analysis, with a fixed search and coding protocol, dual coding with reliability statistics, and a clear statement of how near-direct versus proxy matches were determined, or explicitly state that the novel findings await external confirmation.","section":"App. F; Table 15"}],"minor_comments":[{"comment":"The phrase 'automatically executes a full pipeline' overstates the current experiments: 18 of the 79 tested hypotheses were included manually rather than by the ranking gate (App. C.1), and the user-supplied behavior descriptions in App. C.1 already embed baseline-tendency information that steers hypothesis directions (task 3.baseline in App. A.2). The paper discloses this, but the main text should state it in one sentence so the automation claim is not read as fully hands-off.","section":"Abstract; Sec. 3.1; App. C.1"},{"comment":"Several causal-variable labels are truncated ('Uncertainty of sanctions for aggr...', 'Penalty for misplaced trust', 'Peer purchasing descriptive norms'), and the log10 BF10 axis is truncated at 15, hiding the dynamic range of the most decisive results; the figure should be legible without consulting Table 11.","section":"Fig. 3"},{"comment":"The 26/79 count is presented without a multiplicity calibration. Under a global null with the stated prior and decision thresholds, some fraction of the 79 tests would be expected to reach BF10 >= 3 by chance; a sentence reporting the expected number under the global null (or an equivalent FDR-style computation) would let readers calibrate the headline count.","section":"Sec. 3.1"},{"comment":"The generalization analysis fixes the Stage-2 configurations and holds the Stage-4 reviewer and rubric fixed across subject agents, so the reported Spearman rho reflects, in part, the stability of a shared measurement procedure. The cautious 'may generalize' phrasing is appropriate, but the section should note explicitly that this analysis does not address the construct validity of the scores (see major comment 1).","section":"Sec. 3.2"},{"comment":"The robustness analysis reports 11 decision changes when single evidence classes are removed (and 4 when runs with any null are dropped); given that this is roughly 14% of the 79 hypotheses, the text's characterization that the reported evidence pattern is not heavily influenced by rubric construction should be softened or accompanied by a list of which hypotheses flip class.","section":"App. E.1, Fig. 8"},{"comment":"'According to our search' is not reproducible; the appendix should report the search sources, dates, and keywords, and should provide inter-rater reliability for the consistency coding if the classification is retained.","section":"App. F"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a well-executed systems contribution with unusually careful statistical reporting, and my recommendation is driven by a single load-bearing gap: the Stage-4 dependent variable is never validated against independent human judgment. If the authors add a human-annotation validation (plus, ideally, a different-family reviewer check and a prompt/transcript ablation), I would expect the paper to become acceptable. A few editorial-level points: (i) the abstract's 'first' claim should be checked against the automated behavioral-elicitation tools cited in the paper itself (Petri, Bloom, ALI-Agent); the controlled-experiment distinction is real, but 'first' is stronger than needed; (ii) the 'including some novel ones' phrasing should be toned down unless Appendix F is strengthened; (iii) the GitHub repository is the sole basis for the code/data availability claim and could not be verified from the manuscript; an archived snapshot would be preferable; (iv) the paper would benefit from a sentence calibrating the 26/79 count under multiplicity."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take after reading AEROBAT. The system itself is the real contribution: it takes a target behavior and runs the whole cycle — hypothesis generation, matched controlled experiments, blind review, Bayes factor analysis, report writing — automatically, with code and data shipped. That is new. Prior work like Petri and Bloom does automated elicitation for safety stress tests, and ASS uses LLMs as subjects, but none of them combines arbitrary target behaviors, parametrized environments with multiple realizability, matched controls, and a full statistical pipeline. The authors also did careful engineering: the environment model is thoughtful, the fidelity checks (inverse mapping tasks, config-to-simulation matching) are clever, and the statistical analysis is rigorous, with closed-form Bayes factors, prior sensitivity, and Monte Carlo error estimates. I believe the pipeline works as described.\n\nThe soft spot is the dependent variable. Every reported effect — the 26 findings, the effect sizes, the generalization across subject agents — is computed from the blind reviewer's rubric-based scores (y_hat in Eq. 7). The rubric is generated by GPT-5.1, the same model family that designs the environments and writes the reports. Appendix E.1 shows the five evidence classes are internally consistent (alpha ~0.90), but internal consistency is not criterion validity. The fidelity checks confirm the environments instantiate the manipulated variables; they do not show the scores measure the target construct. If the reviewer latches onto surface textual patterns — e.g., \"strategic,\" \"selective,\" conflict framing — the monotone effects could be artifacts of the scoring procedure rather than properties of the behavior. That is the load-bearing issue for the paper's \"automated discovery\" claim. The stress-test note is right about this.\n\nSeveral smaller concerns: 79 hypotheses tested with no correction for multiple testing; BF10 >= 3 is generous, so some positives are expected by chance. The authors require consistency across domains and groups, which helps, but it isn't a formal control. The manual selection of hypotheses is disclosed and arguably fine for a systems paper, but it biases toward interesting hypotheses. The comparison to prior literature (Appendix F) is post-hoc, author-coded, and mostly \"proxy consistency,\" so it doesn't validate the novel findings. The generalization across subject agents is reassuring but uses the same reviewer, so it doesn't break the circularity.\n\nNone of this kills the paper. The primary claim — that automated behavioral research on AI agents is feasible — is supported. The findings should be framed as candidate discoveries awaiting human confirmation. The fix is straightforward: score a subset of transcripts by independent human raters against the same rubrics, or against their own rubrics, and report agreement. That would also let the authors check whether the rubric itself is construct-valid.\n\nI would send this to peer review. It deserves a serious referee, and the conditional issues are addressable. If the authors add human validation or explicitly reframe the findings as hypotheses, it becomes a solid contribution to the field.","headline":"AEROBAT is a genuine engineering contribution to automating behavioral research on AI agents, but its headline findings are only as strong as its unvalidated LLM-generated behavior scores.","tokens_in":55279,"tokens_out":3618,"would_cite":true,"duration_ms":35630,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Automated pipeline finds 26 behavioral effects in AI agents","keywords":["AI agent behavior","automated behavioral research","multi-agent system","controlled experiments","LLM agents","Bayesian monotone model","hypothesis generation","behavioral assessment"],"falsifier":"Take a sample of the simulation transcripts for several reported findings (for instance, the deception or literal-instruction-following results) and have independent human raters score the same transcripts against the same rubric; if human scores do not track the blind reviewer's scores, or if the reported effect sizes disappear under human scoring, the central claim that the pipeline produces meaningful behavioral findings would be falsified.","tokens_in":54212,"feed_emoji":"🧪","tokens_out":2685,"duration_ms":23052,"temperature":0.7,"pith_summary":"This paper presents AEROBAT, a multi-agent system that takes a user-specified target behavior of an AI agent and automatically runs the whole cycle of behavioral science: generating hypotheses, designing controlled simulation experiments, executing them, scoring the agent's behavior, analyzing results, and writing reports. The authors used it on 12 behaviors, testing 79 hypotheses through 1,240 controlled experiments and 23,512 simulation rounds, and report moderate-to-strong statistical evidence for 26 of them. The central claim is that automated behavioral research can complement manual research and scale it to arbitrary behaviors without hand-built environments for each one.","feed_headline":"Automated pipeline finds 26 behavioral effects in AI agents","feed_subtitle":"An LLM multi-agent system generated, ran, and scored 79 behavioral experiments on agents, finding moderate-to-strong evidence for 26…","key_machinery":"The environment model and four-stage multi-agent pipeline. Environments are parametrized by a domain, abstract environmental variables, and ordinal values, mapped through a configuration with roles, authority, constraints, world updates, and consequence payoffs, with multiple realizations of each variable value. A manager agent gates each stage; the pipeline generates hypotheses, designs matched configurations where only the hypothesized cause varies, renders multi-round simulations in parallel, and has a blind reviewer score behavior against a rubric.","core_discovery":"AEROBAT automatically performs behavioral scientific research on AI agents for an arbitrary target behavior, producing testable hypotheses, matched controlled experiments with varying hypothesized causes, blind behavioral scoring, statistical analysis, and research reports. The authors instantiated the system across 12 social, economic, and operational behaviors and report 26 of 79 tested hypotheses with Bayes factor support and consistent effect sizes across multiple domains, including findings that instruction divergence reduces literal rule-following and that goal conflict increases strategic omission and misleading communication rather than outright lying. They also report that effect sizes generalized across three different subject LLMs and that generated configurations faithfully instantiating environmental variables.","pith_inferences":["If the reported effect sizes replicate across other subject models and scoring rubrics, the pipeline could serve as a standardized probe for mapping 'behavioral phenotypes' of new agent releases without bespoke test design.","The environment model's explicit variable-to-configuration mapping might be turned into an audit tool for AI safety, checking which policy-relevant levers actually change agent behavior in deployment-like settings.","A testable extension would be to run the same hypotheses with human-labeled behavior scores on a sample of transcripts, to see whether the LLM-reviewer scores agree with independent human judgment about the target construct."],"forward_implications":["Behavioral research on AI agents can be conducted at a scale impossible manually, covering an arbitrary target behavior and many candidate causes.","The pipeline's matched-configuration design and multi-domain replication provide a way to attribute behavioral changes to specific environmental variables rather than incidental text or situation.","The 26 supported hypotheses give concrete, testable claims about which environmental factors shape agent behavior, for example that goal conflict increases deceptive omission, instruction divergence reduces literal compliance, and coercive tool access increases strategic aggression.","The system's reports can be used as a first-pass experimental screening tool, letting researchers keep, refine, or discard hypotheses before investing in manual studies."],"supporting_citations":[{"why":"Provides the Bayesian monotone-increment model used for the primary statistical evidence (Bayes factors on ordinal predictors).","marker":"[6]"},{"why":"Source of the multiple-realizability concept the environment model builds on.","marker":"[41]"},{"why":"Prior manual behavioral research on peer influence in LLM agents that AEROBAT contrasts with as not automated.","marker":"[10]"},{"why":"Prior manual behavioral research on sycophancy that AEROBAT contrasts with as not automated.","marker":"[16]"},{"why":"Prior manual benchmark for AI deception behaviors, a comparison point for AEROBAT's automation claims.","marker":"[20]"},{"why":"Prior work using LLMs as simulated subjects, which AEROBAT contrasts with because the research pipeline remains manual.","marker":"[32]"},{"why":"Automated behavioral elicitation tool whose simpler, less structured environments AEROBAT distinguishes from its controlled-experiment design.","marker":"[17]"},{"why":"Automated behavioral evaluation tool AEROBAT contrasts with for lacking controlled experiments around a hypothesized cause.","marker":"[19]"},{"why":"AEROBAT's code and data repository, supporting reproducibility of the reported experiments.","marker":"[25]"},{"why":"Prior work on simulating individuals with LLM agents, relevant to the comparison of automated research pipelines.","marker":"[39]"}],"fun_headline_variants":["AEROBAT automates AI behavioral research, finds 26 effects","Auto pipeline for AI behavior: 79 hypotheses, 26 confirmed","AI agent behavior studies: AEROBAT runs 1,240 experiments","Automated behavioral science on AI agents: 26 evidence-backed effects"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The LLM-generated rubric and the blind reviewer's scores are treated as valid measures of the target behavior, but the scores are not checked against independent human labels for the same simulation transcripts.","fun_headline_variants_meta":{"raw":{"variants":["AEROBAT automates AI behavioral research, finds 26 effects","Auto pipeline for AI behavior: 79 hypotheses, 26 confirmed","AI agent behavior studies: AEROBAT runs 1,240 experiments","Automated behavioral science on AI agents: 26 evidence-backed effects"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00029,"raw_usage":{"total_tokens":1635,"prompt_tokens":821,"completion_tokens":814,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":437,"completion_tokens_details":{"reasoning_tokens":736}},"tokens_in":437,"tokens_out":814,"duration_ms":7232,"temperature":1.0,"reasoning_tokens":736,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:18:14.165964+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a sample of the simulation transcripts for several reported findings (for instance, the deception or literal-instruction-following results) and have independent human raters score the same transcripts against the same rubric; if human scores do not track the blind reviewer's scores, or if the reported effect sizes disappear under human scoring, the central claim that the pipeline produces meaningful behavioral findings would be falsified.","supporting_citations":[{"cited_title":"Psychological predicates","cited_arxiv_id":null,"evidence_quote":"Source of the multiple-realizability concept the environment model builds on."},{"cited_title":"Deceptionbench: A comprehensive benchmark for AI deception behaviors in real-world scenarios","cited_arxiv_id":null,"evidence_quote":"Prior manual benchmark for AI deception behaviors, a comparison point for AEROBAT's automation claims."},{"cited_title":"Automated Social Science: Language Models as Scientist and Subjects","cited_arxiv_id":"2404.11794","evidence_quote":"Prior work using LLMs as simulated subjects, which AEROBAT contrasts with because the research pipeline remains manual."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Automated behavioral elicitation tool whose simpler, less structured environments AEROBAT distinguishes from its controlled-experiment design."},{"cited_title":"Bowman, and Sara Price","cited_arxiv_id":null,"evidence_quote":"Automated behavioral evaluation tool AEROBAT contrasts with for lacking controlled experiments around a hypothesized cause."},{"cited_title":"Aerobat code and data repository","cited_arxiv_id":null,"evidence_quote":"AEROBAT's code and data repository, supporting reproducibility of the reported experiments."}],"review_version":1}