{"id":"d290dc8f-fb2e-4ae2-9765-94603e075247","arxiv_id":"2602.18527","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"JAEGER extends audio-visual LLMs with RGB-D and first-order ambisonics to achieve accurate 3D sound-source grounding and reasoning in simulated environments.","lead":"This paper introduces JAEGER, an audio-visual language model that adds 3D depth and multi-channel spatial audio to locate and reason about sound sources in simulated 3D rooms. It also presents a 61k-sample benchmark and shows that removing spatial audio collapses reasoning accuracy from ~99% to ~44%.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2's 2D baseline is zero-shot, not fine-tuned, and the depth ablation undercuts the full RGB-D necessity claim.","rationale":"The reader's weakest assumption—simulation realism—is plausible but not the first thing to test: even accepting the simulator as a faithful proxy, the reported evidence for the central 3D-vs-2D claim is internally incomplete. The abstract says JAEGER 'consistently surpasses 2D-centric baselines', and §5.2 says 2D AV-LLMs 'fail even after fine-tuning', but the only such baseline is zero-shot, and its below-random scores suggest task-format mismatch, not an inherent modality gap. Meanwhile, Table 5's own depth ablation shows FOA alone with 2D visual features is near ceiling on reasoning, so the necessity claim should be about FOA, not the full RGB-D package. This is a correctness risk in the strength of the claim, independent of the simulator question. The proposed test would settle whether the 3D-superiority claim survives a fair, fine-tuned comparison; if the fine-tuned 2D baseline remains near random, the core method still works but the paper must be reworded to attribute the gain to FOA, not to explicit 3D visual geometry. Secondary issues—inconsistent code/data availability statements and lack of error bars—reinforce the need for conditional acceptance rather than full acceptance.","tokens_in":12981,"tokens_out":11973,"duration_ms":116708,"concrete_test":"Fine-tune two baselines on the SpatialSceneQA D/E training split under JAEGER's exact LoRA recipe: (a) Qwen2.5-Omni with RGB+monaural audio, and (b) JAEGER with the depth encoding removed (RGB+FOA). Evaluate on the identical test split. If (b) still achieves ~99% while (a) remains near random, the claim should be narrowed to 'FOA is necessary; depth is not'; if (a) substantially exceeds random, Table 2's zero-shot comparison is unfair and the headline superiority claim is not supported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—that explicit 3D modeling via RGB-D and FOA is necessary for reliable spatial reasoning—is not established by the reported experiments. First, the only 2D-centric omni baseline, Qwen2.5-Omni, is evaluated zero-shot (§5.2), yet the analysis asserts 2D AV-LLMs 'fail even after fine-tuning'; no fine-tuned 2D baseline appears anywhere in Table 2 or the text. Its reasoning scores (35.8/44.0) are below random (45.6/47.4), consistent with a task-format mismatch rather than a demonstrated modality ceiling. Second, the ablation that is supposed to show '3D necessity' is removal of the FOA encoder (Table 5), which collapses accuracy to 43.8-47.6%. But that only shows directional audio is necessary for a task that is by construction about audio direction. It does not show that RGB-D depth is necessary. In fact, Table 5 shows removing depth alone leaves reasoning accuracy near ceiling (96.9/94.9 with Neural IV; 98.7 with Classical IV in the 2-speaker case). The 'explicit 3D modeling' claim is therefore carried almost entirely by FOA, not by 3D visual geometry. The paper overstates the necessity of the full RGB-D+FOA pipeline.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes JAEGER, an audio-visual large language model that augments Qwen2.5-Omni with RGB-D visual streams and first-order ambisonics (FOA), together with a learned Neural Intensity Vector (Neural IV) for spatial audio. It also introduces SpatialSceneQA, a 61k-sample synthetic benchmark built with SoundSpaces 2.0 and Habitat-Sim, covering single- and overlapping-source DoA estimation, 3D visual grounding, and multi-speaker matching. Experiments report strong performance on the synthetic benchmark, and ablations show that removing the FOA encoder collapses joint-reasoning accuracy to near-random levels, while Neural IV outperforms classical intensity-vector features.","tokens_in":13325,"tokens_out":7355,"duration_ms":69474,"significance":"If the results hold, the paper would provide a useful step toward end-to-end 3D audio-visual grounding: a unified integration of metric depth and directional audio into an existing AV-LLM, a large-scale simulated benchmark with precise 3D annotations, and a learnable spatial audio representation with improved cross-scenario generalization. The scene-based train/val/test splits, held-out loudspeaker assets, and cross-evaluation matrix are good experimental practices. However, the headline claim that explicit 3D modeling is necessary for physical reasoning is broader than the evidence, and several load-bearing comparisons lack appropriate baselines and statistical support.","major_comments":[{"comment":"The only open-source omni baseline, Qwen2.5-Omni, is evaluated zero-shot, but the text and Introduction claim that 2D AV-LLMs 'fail even after fine-tuning.' No fine-tuned monaural/RGB baseline is reported. The below-random accuracy of Qwen2.5-Omni (35.8/44.0 vs random 45.6/47.4) is consistent with a task-format or instruction-following failure, not a demonstrated modality ceiling. Please fine-tune Qwen2.5-Omni (or an equivalent monaural/RGB model) on SpatialSceneQA under the same protocol, or remove the 'even after fine-tuning' claim.","section":"§5.2, Table 2"},{"comment":"The ablation supports the necessity of FOA, not of full RGB-D. Removing the FOA encoder drops 1-/2-speaker reasoning accuracy to 43.8/47.6, near random, but removing depth alone leaves accuracy at 96.9/94.9 (Neural IV) and 99.2/98.7 (Classical IV). Table 4 also shows only modest depth gains (mean 3D IoU 0.29→0.32, median visual offset 0.18→0.16 m). The Conclusion's statement that 'explicit 3D modeling of both visual depth and spatial audio is indispensable' is therefore overstated; the evidence shows directional audio is indispensable and depth provides a small improvement.","section":"§5.3, Table 5"},{"comment":"All metrics are reported as point estimates without variance, confidence intervals, or significance tests. This is particularly important for the Neural IV vs Classical IV comparison, where key margins are small (13.13° vs 16.09° overlap DoA; 99.2 vs 98.6 on 2-speaker reasoning; 94.9 vs 98.7 in the w/o-depth condition). Please report multiple seeds or bootstrap intervals for at least the main comparisons so the reader can assess whether these differences are reliable.","section":"Tables 2–5"},{"comment":"The evaluation is entirely confined to SoundSpaces 2.0/ Habitat-Sim simulations; no real-world recordings or validation are provided. The abstract and conclusion generalize to 'physical environments' and 'physical reasoning tasks' without evidence that simulated FOA/RGB-D renderings capture the difficulty of real acoustic scenes. Please either temper the claims to simulated environments or add a real-world validation set, even a small one.","section":"§5, Abstract/Conclusion"}],"minor_comments":[{"comment":"The abstract states source code, checkpoints, and datasets are available at a URL, while the last line of the full text says they 'will be released upon acceptance.' Please reconcile this inconsistency.","section":"Abstract vs. Section 6 / footnote"},{"comment":"The text refers to 'JAEGER-3D' in one place, but the model is named JAEGER throughout; please make the naming consistent.","section":"§5.2"},{"comment":"There is a stray '61k1.' label in the figure overview; please remove or correct it.","section":"Figure 1"},{"comment":"The reasoning tasks are described as multi-choice among Left/Center/Right, but the random baseline is reported as 45.6/47.4. If there are three visible candidates, random chance should be near 33%; please clarify the choice set or the random baseline construction.","section":"Table 1 / §5.1"},{"comment":"For Neural IV, please clarify whether the 1D-CNN frontend uses shared or separate weights across the omnidirectional f_W and directional f_C channels, and how the latent frame rates are aligned before the element-wise product in Eq. (4).","section":"§4.2 / Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The core method and dataset are potentially useful, but the paper's central 'necessity of explicit 3D modeling' claim is not supported by the experiments as reported. The requested fine-tuned 2D baseline and error bars are essential before acceptance; without them, the headline claims would remain overstated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Rough read: this is a competent, simulation-only paper with a genuinely new benchmark and a useful learned spatial-audio representation. The central claim about the necessity of the full RGB-D+FOA setup does not fully hold up — the ablations show the gains come almost entirely from FOA, not from depth, and the only 2D omni baseline is evaluated zero-shot, not fine-tuned.\n\nWhat is new: first end-to-end AV-LLM trained on RGB-D plus first-order ambisonics; Neural IV as a learnable replacement for classical intensity vectors; and SpatialSceneQA, a 61k-sample synthetic benchmark with DoA, 3D grounding, and multi-speaker matching. The data pipeline is carefully built: scene-based splits, LibriSpeech train/dev/test separation, visible-speaker constraints, and synthetic speaker assets split to test generalization. Cross-evaluation of Classical vs Neural IV is the right way to show robustness, and the FOA ablation (collapse to roughly random without directional audio) is clean evidence that directional audio matters.\n\nSoft spots. The depth ablation (Table 5) undercuts the 'explicit 3D modelling is indispensable' claim. Removing depth leaves reasoning accuracy at 96.9/94.9 with Neural IV and 98.7 with Classical IV in the 2-speaker case. That is a small drop from near ceiling; it does not show depth is necessary for these reasoning tasks. The argument would be stronger if the paper said 'FOA is necessary, depth helps a bit for grounding' rather than 'RGB-D+FOA is indispensable.'\n\nThe only 2D omni baseline, Qwen2.5-Omni, is evaluated zero-shot, yet the text says 2D AV-LLMs 'fail even after fine-tuning.' That claim is not backed by an experiment. Its sub-random scores are more likely a format or modality mismatch than a demonstrated ceiling.\n\nAll metrics are point estimates — no variance, no significance tests. Given the near-perfect scores, it is worth knowing whether the reasoning tasks are too easy or whether the model is robust across scenes.\n\nNo real-world validation. The whole 'physical environments' framing rests on SoundSpaces 2.0 realism. That is fine for a simulation benchmark, but the paper should be more careful in claiming necessity for real environments.\n\nMinor: the availability statement is inconsistent — abstract says available at a URL; full text says released upon acceptance.\n\nWho this is for: people building AV-LLMs or spatial grounding systems, and anyone who wants a synthetic benchmark to test 3D audio-visual reasoning. It deserves a serious referee; the contributions are real, but the claims need to be scaled back to match the evidence. I would ask for a fine-tuned 2D baseline, error bars, and a rewrite of the necessity argument. Yes, I would send it to review rather than desk-reject.","headline":"Solid simulation-only contribution with a new benchmark and a useful learned spatial-audio representation, but the paper overclaims the necessity of RGB-D depth — the ablations point to FOA as the load-bearing input, and the only 2D baseline is zero-shot.","tokens_in":13780,"tokens_out":3620,"would_cite":true,"duration_ms":29823,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Audio-visual language models need explicit 3D perception—depth plus spatial audio—to locate and reason about sound sources; JAEGER reaches 99.2% accuracy on joint reasoning in simulated rooms.","keywords":["audio-visual large language models","3D visual grounding","first-order ambisonics","Neural Intensity Vector","direction-of-arrival estimation","SpatialSceneQA","RGB-D perception","overlapping speaker localization"],"falsifier":"Run the released model, without retraining, on two overlapping speakers in a real furnished room whose positions are known, recorded with a first-order ambisonic microphone and an RGB-D camera, and compare predicted azimuth/elevation and speaker matching to ground truth. If the overlap median angular error is far above the simulated 13.13° or matching accuracy falls toward chance, the claim that explicit 3D modeling plus Neural IV suffices for physical audio-visual reasoning would be falsified.","tokens_in":12923,"feed_emoji":"🎙️","tokens_out":8427,"duration_ms":82502,"temperature":0.7,"pith_summary":"The paper argues that current audio-visual large language models fail at spatial reasoning because they consume 2D RGB images and monaural audio—a dimensionality mismatch that leaves sound-source direction and 3D position ambiguous. JAEGER closes this gap by feeding the model depth-aligned RGB and four-channel first-order ambisonics, and by replacing handcrafted STFT intensity features with a learned 'neural intensity vector.' On a new 61k-sample simulated benchmark, SpatialSceneQA, the model localizes single and overlapping sources with median angular errors of 2.21° and 13.13°, grounds speakers in 3D with 0.32 IoU and 0.16 m center error, and matches audio sources to visible speakers at 99.2% accuracy. Removing the spatial-audio encoder collapses reasoning to near chance, which the authors take as evidence that explicit 3D modeling is indispensable for physical audio-visual reasoning.","feed_headline":"Depth plus spatial audio identifies speakers at 99.2%","feed_subtitle":"A language model that sees depth and hears ambisonic audio can pinpoint overlapping voices and match them to visible speakers.","key_machinery":"The load-bearing mechanism is the Neural Intensity Vector: rather than extracting a fixed STFT-based active intensity vector from the omnidirectional and directional channels of first-order ambisonic audio, JAEGER encodes all four raw channels with a shared 1D CNN, multiplies the omnidirectional latent feature elementwise with each directional latent feature, concatenates the three axes, and maps the result through an MLP into a spatial embedding. This transplants the physics of acoustic intensity—the product of pressure and particle velocity—into a learned latent space, giving direction cues that survive reverberation and overlapping sources. The visual side uses a depth-projected 3D positi","core_discovery":"On its own terms, the contribution is a demonstration that an end-to-end audio-visual LLM can be made 3D-aware without giving up language ability, and that the added modalities carry the task. The authors show that when RGB-D observations and FOA channels are jointly modeled, a single model can estimate direction of arrival in spherical coordinates, regress metric 3D boxes for sound-emitting objects, and resolve which of several visible speakers matches an uttered voice—even when two voices overlap. The headline evidence is the ablation: deleting the FOA encoder sends joint-reasoning accuracy from ~99% to 43.8–47.6%, indistinguishable from random, while deleting depth alone costs a few point","pith_inferences":["Because every reported number comes from simulation, the headline accuracy should be read as a simulation-level result; transferring to real rooms with real microphones and opaque surfaces is an untested leap that the paper's own framing ('simulated physical environments') implicitly concedes.","The near-random FOA ablation shows the benchmark's reasoning tasks are built on directional disambiguation; it demonstrates that monaural audio lacks the necessary information for these questions, not that all audio-visual language understanding requires ambisonics.","The 3D IoU of 0.32 is modest, and removing depth costs only a few accuracy points on reasoning, so the 99% matching scores likely reflect coarse left/center/right localization plus strong audio direction rather than fine metric grounding; finer spatial questions might expose a smaller audio-visual alignment than the headline suggests.","The Neural IV recipe—learned products of omnidirectional and directional latent channels—is a generic inductive bias that could transfer to other microphone geometries or to self-supervised pretraining, since it does not depend on STFT binning or array-specific filters."],"forward_implications":["Adding FOA spatial audio to an AV-LLM changes joint speaker matching from chance to near-solved in simulation: accuracy goes from 43.8–47.6% without the FOA encoder to 98.6–99.5% with it.","The learned Neural IV beats the classical intensity vector under overlapping sources (13.13° vs 16.09° median angular error) and degrades less when training and test source configurations are mismatched, suggesting learnable spatial representations are more robust than handcrafted features.","Depth injection improves 3D grounding moderately—mean IoU 0.29→0.32 and median center offset 0.18→0.16 m—and also lifts reasoning accuracy, so metric geometry helps but is not the main driver of the reasoning gains.","A single model can natively output spherical directions, metric 3D boxes, and speaker identities after low-rank fine-tuning, removing the need for separate localization modules and signal-processing front-ends.","SpatialSceneQA provides 61k synchronized RGB-D plus 4-channel FOA samples with exact 3D annotations across held-out scenes and speaker meshes, making it a reusable training and evaluation resource for joint audio-visual grounding."],"fun_headline_variants":["3D audio-visual LLM: depth + ambisonics pinpoint speakers at 99%","Without spatial audio, LLM speaker ID crashes to ~45%","JAEGER: LLM sees depth, hears ambisonics, grounds 3D voices","Depth+FOA gives LLMs 3D scene understanding for speaker ID","2D audio-visual LLMs fail: 3D modeling is key, JAEGER shows"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the simulated sound propagation, ambisonic microphone rendering, and depth images faithfully capture the difficult parts of real rooms; all headline results come from simulation, with no real-world validation.","fun_headline_variants_meta":{"raw":{"variants":["3D audio-visual LLM: depth + ambisonics pinpoint speakers at 99%","Without spatial audio, LLM speaker ID crashes to ~45%","JAEGER: LLM sees depth, hears ambisonics, grounds 3D voices","Depth+FOA gives LLMs 3D scene understanding for speaker ID","2D audio-visual LLMs fail: 3D modeling is key, JAEGER shows"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001106,"raw_usage":{"total_tokens":4451,"prompt_tokens":753,"completion_tokens":3698,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":497,"completion_tokens_details":{"reasoning_tokens":3602}},"tokens_in":497,"tokens_out":3698,"duration_ms":28510,"temperature":1.0,"reasoning_tokens":3602,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T22:01:55.679416+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the released model, without retraining, on two overlapping speakers in a real furnished room whose positions are known, recorded with a first-order ambisonic microphone and an RGB-D camera, and compare predicted azimuth/elevation and speaker matching to ground truth. If the overlap median angular error is far above the simulated 13.13° or matching accuracy falls toward chance, the claim that explicit 3D modeling plus Neural IV suffices for physical audio-visual reasoning would be falsified.","supporting_citations":[],"review_version":1}