{"id":"1d67fb78-0ce3-45d6-871e-3df480e77b47","arxiv_id":"2505.22809","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Large audio-language models can act as overhearing agents that listen to human conversation and provide useful suggestions, as shown in a Dungeons & Dragons study where audio cues helped distinguish in-character from out-of-character speech.","lead":"This paper introduces \"overhearing agents\": AI assistants that listen to human conversations and send suggestions without joining the conversation. Tests on 41.5 hours of Dungeons & Dragons play show that large audio-language models can use tone, pauses, and speaker cues to help the game master, while smaller models and text-only systems struggle.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The audio-vs-text comparison for GPT-4o/GPT-4o-mini is confounded: audio and text runs use different model checkpoints (Appendix B, Table 5), so the reported audio advantage may be a checkpoint effect rather than evidence for implicit audio-cue use.","rationale":"The reader's weakest assumption concerns the 300-second matching window and the reliability of gold suggestions. That is a legitimate evaluation concern, but it is not the most load-bearing issue for the central claim. The paper's abstract claims that large audio-language models have an 'emergent ability' to use implicit audio cues. The quantitative case for that claim rests on audio outperforming text for GPT-4o and GPT-4o-mini. Appendix B, Table 5 shows those two conditions used different model checkpoints, so the comparison is confounded: a newer or differently aligned checkpoint could explain the gap even if the models use no audio-specific cues at all. The open-weight models, which do hold the checkpoint fixed, provide no positive evidence of an audio advantage. Therefore the central claim is not established by the reported experiments. I keep the verdict at CONDITIONAL rather than REJECT because the flaw is fixable with a checkpoint-matched rerun, the paper releases its code and tools, and the qualitative examples may still point in the right direction. The acceptance condition should explicitly require a controlled modality comparison on the same checkpoint, with uncertainty quantification on the audio-vs-text difference.","tokens_in":18715,"tokens_out":8589,"duration_ms":84292,"concrete_test":"Re-run the GPT-4o and GPT-4o-mini audio/text comparisons using the same checkpoint in both conditions—for example, use gpt-4o-realtime-preview-2024-12-17 for both audio and text (feeding Whisper transcripts as the text input), keeping prompts, tools, and temperature identical. If the audio F1 no longer exceeds text F1 by a clear margin, the 'implicit audio cues' claim is unsupported; report bootstrapped 95% confidence intervals on the difference.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that large audio-language models can perform overhearing-agent tasks 'using implicit audio cues.' The primary quantitative support is the audio-vs-text F1 gap for GPT-4o (0.369 vs 0.236) and GPT-4o-mini (0.159 vs 0.042) in Figure 3 / Table 4. This comparison is not modality-controlled. Appendix B, Table 5 lists the audio condition as gpt-4o-realtime-preview-2024-12-17 / gpt-4o-mini-realtime-preview-2024-12-17 and the text condition as gpt-4o-2024-11-20 / gpt-4o-mini-2024-07-18. These are different checkpoints, likely with different post-training, tool-calling behavior, and decoding characteristics, so the observed audio advantage could arise without any use of prosody, speaker identity, or hesitation cues. For the open-weight models, weights are matched across modalities, but none shows an audio advantage: Ultravox performs better with text, and Qwen2.5/Phi-4 are near zero in both modalities. Thus the only positive quantitative evidence for the headline claim comes from the confounded OpenAI comparison. Table 2 is a suggestive qualitative example, but it does not establish the claim. The reader's concern about the 300-second window and gold-label reliability is valid, but the checkpoint confound is more directly load-bearing because it undermines the modality comparison itself.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces \"overhearing agents,\" an interaction paradigm in which an LLM-based agent passively listens to human-to-human conversation and performs background tasks via tool calls, rather than conversing directly with the user. The paradigm is instantiated as a Dungeon Master's assistant for live D&D games, with three tasks: game-data retrieval, NPC stage direction, and improvised NPC generation. The authors recorded 41.5 hours of gameplay, ran five audio-language models in both audio and text-transcription conditions, and evaluated suggestions through stopwatch-based recall annotation and post-hoc human precision annotation. They report F1 scores per model and task, speed measurements, and ablations (removing reasoning, adding transcription). The headline claim is that some large audio-language models have an emergent ability to use implicit audio cues, such as speaker identity, tone, and hesitation, to perform these overhearing tasks. The paper also releases open-source libraries and project code.","tokens_in":19053,"tokens_out":4186,"duration_ms":45995,"significance":"The overhearing-agent paradigm is a genuinely new framing for LLM agent interaction, and the D&D setting is a rich, naturalistic test bed. The study has real strengths: a large novel dataset, human evaluation with two complementary annotation mechanisms (stopwatch recall and post-hoc precision), a moderate inter-annotator agreement of alpha=0.67, multiple model families and sizes, and a public code and library release. If the audio-cue claim were solidly supported, this would be a notable result for multimodal agents and for understanding how audio signals beyond words can be exploited by LLMs. However, as presented, the central evidence for the audio advantage is not yet convincing because the audio and text conditions for the two OpenAI models use different model checkpoints, and the open-weight models, which are checkpoint-matched, show no audio advantage. The evaluation also lacks confidence intervals and significance tests, and the 300-second matching window is loose enough to require sensitivity analysis.","major_comments":[{"comment":"The audio-vs-text comparison for GPT-4o and GPT-4o-mini, which is the primary quantitative support for the \"implicit audio cues\" claim, is confounded by model checkpoint differences. The audio condition uses gpt-4o-realtime-preview-2024-12-17 and gpt-4o-mini-realtime-preview-2024-12-17, while the text condition uses gpt-4o-2024-11-20 and gpt-4o-mini-2024-07-18. These are different model versions with potentially different post-training, tool-calling behavior, and decoding characteristics, so the F1 gaps in Figure 3 (GPT-4o: 0.369 audio vs 0.236 text; GPT-4o-mini: 0.159 vs 0.042) cannot be attributed to audio modality alone. This matters because the open-weight models, which do use identical weights across modalities, show either no audio advantage (Ultravox: text 0.343 vs audio 0.132) or near-zero performance in both modalities (Qwen2.5, Phi-4). The only positive quantitative evidence for the headline claim therefore rests on the confounded comparison. The authors should either run a checkpoint-matched comparison (for example, using a model that accepts both audio and text inputs under identical weights, or matching the text model version to the realtime-preview version if the API permits), or provide a separate controlled experiment that manipulates audio cues (e.g., same transcript with different prosody or speaker diarization) to isolate their effect, and should revise the abstract's claim if such evidence is not available.","section":"Appendix B, Table 5"},{"comment":"The F1 scores are reported as point estimates without confidence intervals or significance tests. Given that the dataset is a single continuous game with 14,939 turns and 940 gold suggestions, the differences between conditions may be within sampling noise; for example, Ultravox-text (0.343) and GPT-4o-audio (0.369) are close, and several small-model scores are near zero. The paper's qualitative conclusions about model ordering and about which models \"significantly\" benefit from audio would be substantially stronger with bootstrap confidence intervals or an appropriate paired significance test across sessions or time intervals. Without these, the reader cannot tell whether the reported F1 differences are reliable, which is load-bearing for the central claim.","section":"Section 4, Figure 3"},{"comment":"The 300-second matching window for a suggestion to be considered correct is very generous relative to both the 10-second input intervals and the real-time nature of the assistive task. Over a 41.5-hour corpus, a model that periodically guesses common entities or repeatedly emits suggestions could match gold events almost by chance, especially for the Generate NPCs task. The paper should report how F1 changes as a function of the matching window (e.g., 30s, 60s, 120s, 300s) and should present precision and recall separately for each task and model, rather than only aggregate F1 and one task-specific precision table. In addition, Section 3.3 is ambiguous about whether the 940 gold suggestions come from the DM's stopwatch events or from filtered positive post-hoc annotations; this distinction should be clarified because it directly affects how recall is computed.","section":"Section 4, evaluation metrics"}],"minor_comments":[{"comment":"The title contains a spacing typo: \"T owards\" should be \"Towards.\"","section":"Title page"},{"comment":"The phrase \"presentoverhearingAI agents\" is missing a space; it should be \"present overhearing AI agents.\"","section":"Abstract"},{"comment":"The sentence \"the agent in which the agent converses directly with the user\" repeats \"the agent\" and is grammatically broken; it should be \"in which the agent converses directly with the user.\"","section":"Section 5"},{"comment":"\"audio-langauge models\" is a typo for \"audio-language models.\"","section":"Section 5"},{"comment":"The heading \"T asks\" should read \"Tasks.\"","section":"Section 3.1"},{"comment":"The y-axis label \"Times Realtime\" should be \"Times real-time\" and the measurement basis (relative to what reference for open-weight models) should be stated more precisely in the caption.","section":"Figure 4"},{"comment":"The table caption refers to errors \"marked in red,\" but the reproduced table does not show color; please mark the erroneous span or otherwise indicate it clearly in print.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The paper introduces a worthwhile paradigm and the empirical infrastructure is substantial, but the headline claim about implicit audio cues currently rests on a checkpoint-confounded comparison. This is fixable in principle by rerunning the text condition with a checkpoint-matched model or by adding a controlled audio-cue manipulation; short of that, the abstract and conclusions need to be narrowed to the claim that audio-language models can perform overhearing-agent tasks, without attributing the advantage specifically to audio cues. The evaluation-metric concerns (window sensitivity, missing confidence intervals) also need addressing before the quantitative claims can be considered reliable. I would not reject the paper: the paradigm, dataset, annotation effort, and code release have independent value, and the technical issue is local rather than a failure of the underlying idea."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things before reading: this paper introduces a real new interaction paradigm, and its central quantitative claim rests on a comparison that doesn't control for model version.\n\nWhat's actually new: 'overhearing agents' are not a chatbot the user talks to, but a passive listener that calls tools in the background while humans talk. Applied to a D&D DM assistant, that's a natural and well-executed testbed. The authors recorded 41.5 hours of gameplay, built three assistive tasks, evaluated five audio-language models against transcription-based text pipelines, and did a human eval with a stopwatch-recall/post-hoc-precision design. The annotation interface, Krippendorff's alpha of 0.67, and the failure-mode analysis are all solid. They also release the code and two Python libraries, which is real infrastructure for future work. On its own terms, the GPT-4o audio run beats the text run (0.369 vs 0.236 F1; mini 0.159 vs 0.042), and Table 2 gives a nice qualitative example of prosody disambiguating in-character speech.\n\nBut the headline claim that these models use 'implicit audio cues' is not actually supported by the quantitative evidence. Table 5 shows the audio runs used gpt-4o-realtime-preview-2024-12-17 and gpt-4o-mini-realtime-preview-2024-12-17, while the text runs used gpt-4o-2024-11-20 and gpt-4o-mini-2024-07-18. Those are different checkpoints with likely different post-training and tool-calling behavior. The open-weight models, where the weights are matched across modalities, show no audio advantage at all—Ultravox does better with text. So the only positive result comes from a confounded comparison, and the claim should at minimum be softened or defended.\n\nSecondary concerns: the 300-second matching window is generous and has no sensitivity analysis; F1 numbers come without confidence intervals or significance tests; gold labels are subjective (though the annotation design is reasonable); and the dataset isn't released, only code. These are minor-to-moderate, not fatal.\n\nNet: this is an honest, useful study with a novel framing and good infrastructure. The checkpoint confound is the one thing I'd want a referee to force them to address, either by re-running the text condition on the realtime checkpoint or by explicitly arguing the two are equivalent for this task. The paradigm deserves follow-up even if the strongest claim doesn't fully hold.\n\nMy recommendation: send it to peer review. It should be published after the confound is fixed or acknowledged. I'd bring it to reading group.","headline":"A genuinely new paradigm and a serious empirical study, but the headline audio-vs-text claim is confounded by checkpoint differences in the only models that show the effect.","tokens_in":19548,"tokens_out":3635,"would_cite":true,"duration_ms":39590,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large audio-language models can act as 'overhearing agents' that listen to human-to-human conversation and make background suggestions, with the biggest model using tone, pauses, and speaker identity that transcripts cannot capture.","keywords":["overhearing agents","audio-language models","multimodal LLM agents","tool calling","Dungeons & Dragons","implicit audio cues","human evaluation","real-time AI assistant"],"falsifier":"Take the recorded gameplay audio and strip out exactly the cues the paper says matter, for example flattening pitch and speaking rate or scrambling which voice is the DM's while keeping the words identical, and rerun GPT-4o on the audio-input setting; if its F1 does not drop materially below 0.369, the audio-channel advantage is not being driven by tone, pause, and speaker-identity cues as claimed.","tokens_in":18518,"feed_emoji":"🎲","tokens_out":6314,"duration_ms":61574,"temperature":0.7,"pith_summary":"This paper introduces \"overhearing agents\": AI assistants that never speak to their user but listen in on human-to-human conversation and hand helpful suggestions to the background, through tool calls rather than dialogue. To test whether such agents work, the authors built one that assists a Dungeon Master during live D&D games, with three background tasks (retrieving game rules, managing NPC portraits on a virtual stage, and improvising new NPCs), and ran it on 41.5 hours of recorded play. The central claim is that some large audio-language models have an emergent ability for this task: the strongest model, GPT-4o, reached an F1 of 0.369 with audio input versus 0.236 when given only a transcript, and the gap is credited to implicit audio cues such as the DM's tone, speaking rate, and hesitation. A reader should care because, if true, the result shows that passive, non-intrusive assistants for meetings, calendars, and other group settings are within reach of existing models, and that the audio channel carries information that text transcription throws away.","feed_headline":"Audio AI overhears D&D games and aids the Dungeon Master","feed_subtitle":"The best model nearly doubles in helpfulness when it hears tone and pauses, not just words.","key_machinery":"The load-bearing machinery is a real-time tool-calling loop: gameplay audio is fed to a multimodal model in 10-second intervals appended to one running conversation (capped at 15 minutes), and each round the model must write chain-of-thought reasoning and then either call a tool or output \"None.\" The three tasks are exposed as functions (search D&D data, manage NPCs on a virtual stage, generate an improvised NPC), so the model can only \"speak\" by making suggestions through those calls. What makes the paradigm work is the reasoning step: the model formulates an internal belief about the conversational goal before acting, and the paper shows that removing that step costs most of the performance. The audio itself carries the decisive cues, speaker identity and in-character voice for the stage-director task, pauses and filler words for NPC generation, which is why transcript-based systems hit a ceiling that the audio models pass.","core_discovery":"On the paper's own terms, the discovery is that overhearing-agent behavior, inferring what a group of people needs and acting on it without joining the conversation, is not something that must be engineered task by task; it emerges in sufficiently large audio-language models. The paper argues this is specifically an audio-channel ability: a model that hears the game can tell when the Dungeon Master is speaking in character versus out of character, can notice pauses and fillers that signal the DM is improvising, and can track long-term conversational goals across 15-minute windows of context. Three lines of evidence carry the claim: the audio-over-text performance gap for the largest models, the collapse (more than 70% average F1 loss) when chain-of-thought reasoning is removed, and the failure-mode analysis showing that text-only pipelines either hallucinate NPC dialogue or slip into a \"conversational default\" of replying to the players directly. Smaller models and an ASR-pretrained audio model do not show the ability, which the paper attributes to instruction-tuning and pretraining objectives that do not prepare models for the overhearing role.","pith_inferences":["Editorial inference: the 300-second matching window is generous enough that near-duplicate suggestions minutes apart count as hits; tightening the window to, say, 60 seconds would likely compress the gap between models and is the single most direct robustness check a skeptic could run.","Editorial inference: the audio advantage should be portable to other group settings with high cognitive load, such as negotiation tables, medical consultations, and live interviews, and a cheap test would be a Wizard-of-Oz calendar agent that only ever acts when conversationally appropriate, mirroring the DM-assistant design.","Editorial inference: the \"conversational default\" failure of text models suggests that instruction-tuned models carry a strong prior toward replying; an overhearing agent may need its output channel (tool calls only, never visible text) to be enforced at the decoding level, not just the prompt level."],"forward_implications":["If the claim holds, transcription-based pipelines are not just a practical shortcut but a ceiling: any task that depends on who said something, how they said it, or when they hesitated will need models that hear the audio itself.","Real-time overhearing agents must budget for chain-of-thought reasoning, since removing it costs more than 70% of F1 on average; the paper's speed numbers show the two can overlap, but only within a narrow window.","The paradigm generalizes by swapping tools: the same passive-listening model, given calendar or meeting tools instead of D&D tools, is claimed to direct the same emergent reasoning toward scheduling and meeting support.","Smaller audio models as they exist today are not usable for this task, they degenerate into repeated suggestions or silence, so the paper's implication is that distillation, not just scaling, is the next lever."],"supporting_citations":[{"why":"Supplies the theoretical frame that overhearers understand less than addressees, which motivates why the overhearing task is harder than conversation.","marker":"Schober & Clark, 1989"},{"why":"Origin of the \"overhearing agents\" concept that the paper extends from multi-agent systems to LLMs serving human users.","marker":"Busetta et al., 2001"},{"why":"The closest prior work: a text-transcript LLM assistant for tabletop game masters whose unimodal approach the paper's audio pipeline must beat.","marker":"Kelly et al., 2023"},{"why":"The D&D Basic Rules sourcebook that defines the entity space the Game Data Retrieval tools search.","marker":"Crawford et al., 2018"},{"why":"Whisper supplies the transcription baseline and also the encoder used by Ultravox, whose ASR-only pretraining is cited as the reason audio does not help it.","marker":"Radford et al., 2023"},{"why":"Documents Ultravox's KL-to-transcript training objective, the evidence for the paper's claim that ASR-pretrained audio models cannot use cues beyond the words.","marker":"Fixie AI, 2025"},{"why":"ReAct-style reasoning is the prompting mechanism whose removal costs most of the performance, supporting the claim that overhearing relies on inference-time reasoning.","marker":"Yao et al., 2023"}],"fun_headline_variants":["Overhearing AI aids D&D Dungeon Masters silently","Audio AI overhears D&D, boosts Dungeon Masters","Overhearing emerges in large AIs, aiding D&D DMs","Tone and pauses teach AI to help D&D Dungeon Masters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation counts a model suggestion as correct when it falls within 300 seconds of a matching human-annotated gold suggestion, and those gold suggestions come from the Dungeon Master's own stopwatch notes plus players' subjective ratings of helpfulness, so the whole \"emergent ability\" conclusion rests on that window and those human judgments being trustworthy.","fun_headline_variants_meta":{"raw":{"variants":["Overhearing AI aids D&D Dungeon Masters silently","Audio AI overhears D&D, boosts Dungeon Masters","Overhearing emerges in large AIs, aiding D&D DMs","Tone and pauses teach AI to help D&D Dungeon Masters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00161,"raw_usage":{"total_tokens":6405,"prompt_tokens":936,"completion_tokens":5469,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":5397}},"tokens_in":552,"tokens_out":5469,"duration_ms":40439,"temperature":1.0,"reasoning_tokens":5397,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:59:38.435936+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the recorded gameplay audio and strip out exactly the cues the paper says matter, for example flattening pitch and speaking rate or scrambling which voice is the DM's while keeping the words identical, and rerun GPT-4o on the audio-input setting; if its F1 does not drop materially below 0.369, the audio-channel advantage is not being driven by tone, pause, and speaker-identity cues as claimed.","supporting_citations":[{"cited_title":"Extending Multi -agent Cooperation by Overhearing","cited_arxiv_id":null,"evidence_quote":"Origin of the \"overhearing agents\" concept that the paper extends from multi-agent systems to LLMs serving human users."},{"cited_title":"D&D Basic Rules","cited_arxiv_id":null,"evidence_quote":"The D&D Basic Rules sourcebook that defines the entity space the Game Data Retrieval tools search."},{"cited_title":"Robust speech recognition via large-scale weak supervision","cited_arxiv_id":null,"evidence_quote":"Whisper supplies the transcription baseline and also the encoder used by Ultravox, whose ASR-only pretraining is cited as the reason audio does not help it."},{"cited_title":"Introducing ultravox v0.5: Taking the lead in speech understanding, 2025","cited_arxiv_id":null,"evidence_quote":"Documents Ultravox's KL-to-transcript training objective, the evidence for the paper's claim that ASR-pretrained audio models cannot use cues beyond the words."}],"review_version":1}