{"id":"8bc59d30-338a-466f-bb13-56cf50cdc586","arxiv_id":"2504.15280","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"All-Angles Bench, a 2,132-question benchmark across 90 real scenes, shows current MLLMs score around 60% on multi-view understanding while humans score 82%, with the largest gaps in camera pose estimation and cross-view correspondence.","lead":"This paper introduces All-Angles Bench, a human-annotated benchmark that tests how well multimodal AI models understand the same scene from multiple camera views. It finds that current models, including GPT-4o and Gemini, are far behind humans on these multi-view reasoning tasks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Human-baseline reliability is the weak link: only two annotators split 250 questions, no agreement metric, and the main text contradicts the supplement on the protocol.","rationale":"The central claim is that current MLLMs remain far from human-level multi-view understanding, quantified by a 21-point gap on a 250-question subset. The most load-bearing assumption is that the human baseline and the ground-truth answers against which models are scored are both reliable. The paper provides only two human evaluators who appear to have answered disjoint halves (Supplement §8.1) despite the main text implying each answered every question; no inter-annotator agreement is reported for either the 250-question subset or the full benchmark. This is a genuine weakness because the exact magnitude of the human-MLLM gap depends on the reproducibility of the 82.0 figure. However, the gap is so large (82 vs. ~60, and even larger on the full benchmark) that even a substantial downward revision of the human baseline would not overturn the qualitative conclusion that models lag humans. The reader's conditional verdict is therefore appropriate: the benchmark is valuable and the main finding is likely correct, but the human-baseline evidence must be strengthened before the precise claim is fully accepted. The proposed concrete test directly addresses this by measuring inter-annotator agreement and the reproducibility of the human score on the same subset.","tokens_in":31955,"tokens_out":9388,"duration_ms":86142,"concrete_test":"Recruit five new PhD annotators who have not seen the benchmark and have each independently answer all 250 questions in the tiny subset (not split). Compute (i) Fleiss' kappa among the five, (ii) each annotator's accuracy against the published ground truth, and (iii) the pooled human accuracy. If the pooled accuracy is within about ±5 points of 82.0 and kappa exceeds 0.7, the human baseline is reproducible and the concern is resolved. If kappa is low or the pooled accuracy differs by more than 10 points, the claimed human-MLLM gap is not stable and the central conclusion would need to be re-quantified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline gap (human 82.0 vs. best MLLM ~60.8) rests on a 250-question subset. The main text (§3.1) states that human annotators 'each of whom independently answers every question,' but Supplement §8.1 says each of the two evaluators was assigned 125 questions, meaning the halves are disjoint. This direct contradiction makes the exact protocol unclear and, with no overlap, inter-annotator agreement cannot be computed. Per-task human accuracies (e.g., 88.9% on camera pose estimation) are based on roughly 20-30 questions, so their confidence intervals are wide. More fundamentally, the ground-truth answers for relative distance, relative direction, and object manipulation are the judgments of eight PhD annotators, with no inter-annotator agreement reported for the full 2,132 questions either. If a nontrivial fraction of questions is ambiguous, both human and model accuracies are measured against a subjective standard, so the precise magnitude of the human-MLLM gap is not established. The gap is large enough to survive modest noise, but the specific claim of 82.0% human accuracy is not reproducible as reported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces All-Angles Bench, a benchmark of over 2,100 human-annotated multi-view question-answer pairs across 90 real-world scenes, spanning six tasks: counting, attribute identification, relative distance, relative direction, object manipulation, and camera pose estimation. The authors evaluate 27 closed- and open-source MLLMs with greedy decoding and an LLM-based answer extractor, and compare against a human baseline on a 250-question subset. They report a substantial human–model gap, identify cross-view correspondence and camera pose estimation as the main failure modes, and show that chain-of-thought prompting yields limited and model-dependent improvements. The central claim is that current MLLMs remain far from human-level multi-view understanding.","tokens_in":32152,"tokens_out":4362,"duration_ms":40361,"significance":"If the benchmark is reliable, it fills a clear gap: most existing evaluations of spatial reasoning in MLLMs are single-view or temporal, whereas All-Angles Bench directly tests geometric correspondence and cross-view consistency. The public release, the breadth of models evaluated, the paired-question consistency design, and the careful annotation effort (eight PhD annotators, human refinement of LLM-generated questions, cross-checking) are concrete strengths. The identification of camera pose estimation as a systemic weakness, with many models below random guessing, is a valuable and falsifiable finding for the embodied-AI community. However, the quantitative headline, particularly the 82.0% human vs. roughly 60% best-MLLM gap, rests on a 250-question human baseline whose exact protocol is contradictory in the manuscript and for which no agreement or uncertainty measures are reported.","major_comments":[{"comment":"The human evaluation protocol is described inconsistently. §3.1 states that human annotators 'each of whom independently answers every question,' while Supplement §8.1 states that each of the two evaluators was assigned 125 questions, i.e., disjoint halves with no overlap. If the latter is correct, the per-task human accuracies in Table 1 are computed on very small subsets and inter-annotator agreement cannot be assessed at all. This is load-bearing for the headline claim of a 82.0 vs. ~60.8 human–model gap. Please clarify the exact protocol, provide the raw per-question human answers, and if overlap exists, report a chance-corrected agreement statistic. If the disjoint design is kept, the human baseline should be re-estimated with confidence intervals and the text should be corrected.","section":"§3.1 vs. Supplement §8.1"},{"comment":"All human and model accuracies on the 250-question subset are reported as point estimates without confidence intervals or significance tests. Based on the task distribution in Figure 4, per-task sample sizes are roughly 20–60 questions; for example, camera pose estimation has about 18 questions in this subset. Consequently, a difference of a few percentage points between models, and even the ordering of some models, may reflect sampling noise. The human–model gap is large enough to survive modest noise for several tasks, but the precise magnitudes (e.g., 82.0% vs. 60.8% on average, or 88.9% vs. 16.7% on camera pose) are not statistically supported as reported. Please provide binomial confidence intervals or a significance test for the key comparisons, especially for the human baseline and the top-performing models.","section":"Table 1 and Table 2 (250-question subset)"},{"comment":"The ground-truth answers for relative distance, relative direction, and object manipulation are annotator judgments, and the full 2,132-question set relies on eight PhD annotators with cross-checking but no reported inter-annotator agreement number. The supplement states that disagreements were resolved through group discussion, but the frequency and nature of disagreements are not quantified. Because the entire evaluation compares model accuracy against these subjective labels, an ambiguous-question analysis (e.g., Cohen's kappa on a shared subset, or annotation confidence scores) is required to rule out the possibility that annotation noise inflates the measured human–model gap. This is a correctness-risk concern, not an assertion that the labels are wrong; the gap is probably large enough to survive some noise, but the reported exact values need this support.","section":"Supplement §7.3 (annotation reliability)"}],"minor_comments":[{"comment":"The text reads 'we hiredeight Ph.D. students' — a typo for 'hired eight'.","section":"Supplement §7.3"},{"comment":"The model name 'LLaV A-Onevision' appears in Table 1 and elsewhere; the standard spelling is 'LLaVA-OneVision', which should be used consistently.","section":"Abstract and Table 1"},{"comment":"The phrase 'benchmark on27 representative MLLMs' is missing a space between 'on' and '27'.","section":"Abstract"},{"comment":"The phrase 'top-bottom view' should be 'top-down view' in the camera pose estimation task description and figures.","section":"Figure 1 and Section 2"},{"comment":"The plot labeled 'Averaged IC' would be easier to interpret if the caption stated explicitly whether the average is a macro-average over the six tasks or a micro-average over questions.","section":"Figure 5"}],"recommendation":"major_revision","confidential_remarks":"The benchmark is potentially a useful contribution and the central finding of a large human–model gap is likely robust in direction. However, the human baseline is the linchpin of the paper's quantitative headline, and the contradictions and missing reliability metrics in the current manuscript are not purely cosmetic. I would ask the authors to fix the protocol description, provide raw human answers and an agreement measure, and add confidence intervals for the main comparisons. If these are provided satisfactorily, the paper could be suitable for acceptance. I do not see a fundamental circularity problem: although GPT-4o generated initial questions, the human refinement and removal of ambiguous items substantially breaks the generation–evaluation loop."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The punchline: All-Angles Bench is a genuinely useful new benchmark, and the central finding — current MLLMs are far behind humans at multi-view understanding — is almost certainly right. The weak link is the human baseline: it rests on a 250-question subset, two annotators, no agreement metric, and the main text and supplement contradict each other about how those annotators were split.\n\nWhat is actually new and good: the benchmark uses 4–5 synchronized real-world views per scene from Ego-Exo4D and EgoHumans, with six tasks that target geometric correspondence rather than single-view or temporal reasoning. The paired-question protocol, where 85.3% of questions have a rephrased counterpart, is a smart way to test whether models truly understand multi-view structure or just answer isolated queries. The annotation pipeline is serious: eight PhD annotators, cross-checking, group discussion for disagreements, and 300 person-hours of work. The evaluation itself is careful — greedy decoding, answer extraction via Qwen2.5-32B, and a 250-question human subset with per-task reporting. The failure analyses (partial-visibility counting, camera pose estimation, the visualization of GPT-4o and Gemini-2.0-Flash misaligning camera coordinates) are concrete and informative.\n\nThe sore spots are statistical and concentrated in the human baseline. The stress-test note is accurate: Section 3.1 says human evaluators \"each of whom independently answers every question,\" while Supplement Section 8.1 says each evaluator was assigned 125 questions, meaning the two halves are disjoint. That is a direct contradiction, and with no overlap, inter-annotator agreement cannot be computed. Per-task human accuracies like 88.9% on camera pose estimation are based on roughly 20–30 questions, so their confidence intervals are wide. The ground-truth answers for relative distance, relative direction, and object manipulation are the judgments of eight annotators, with no agreement reported for the full 2,132 questions either. None of this overturns the main conclusion — the gap between 82.0% human and roughly 60% for the best MLLM is large enough to survive substantial annotation noise — but the exact magnitude is not reproducible as stated, and the missing error bars matter for anyone comparing future models against this benchmark.\n\nThe circularity concern is minor: GPT-4o generated initial questions but humans revised and validated them, which breaks the loop. Self-citation is not a problem here; the benchmark is the contribution, and the dataset is public.\n\nThis paper deserves a serious referee, not a desk reject. I would send it to peer review with the request that the authors reconcile the human-evaluation protocol, add confidence intervals or agreement statistics, and either enlarge the human sample or soften the precise claims about the human–MLLM gap. For the embodied AI and MLLM evaluation community, this is a useful instrument either way.","headline":"A genuinely useful multi-view benchmark with a compelling main result, but the human-baseline statistics are too shaky to pin down the exact gap as reported.","tokens_in":32717,"tokens_out":1703,"would_cite":true,"duration_ms":16801,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces All-Angles Bench, a 2,132-question multi-view benchmark, and shows that leading multimodal models—InternVL2.5-38B, Qwen2.5-VL-72B, Gemini-2.0-Flash—score 55–61 on a human-evaluated subset where humans average 82, with…","keywords":["multi-view understanding","multimodal large language models","benchmark","camera pose estimation","cross-view correspondence","geometric reasoning","paired questions","evaluation"],"falsifier":"Run a fresh panel of at least five independent annotators on a randomly selected 250-question subset, computing inter-annotator agreement; if their average accuracy falls more than 10 points short of the paper's 82.0 human baseline, or if any MLLM reaches within 5 points of the new baseline on a held-out set, the central gap claim would be overturned.","tokens_in":31768,"feed_emoji":"🎥","tokens_out":8792,"duration_ms":72881,"temperature":0.7,"pith_summary":"Current multimodal large language models (MLLMs) do not understand a 3D scene from multiple camera views the way people do, and the paper's aim is to make that gap measurable and precise. It builds All-Angles Bench, a benchmark of 2,132 human-annotated multiple-choice questions over 90 real-world scenes, covering counting, attribute identification, relative distance, relative direction, object manipulation, and camera pose estimation. On a 250-question subset that humans answered at 82.0% accuracy, the best tested models—InternVL2.5-38B, Ovis2-16B, Qwen2.5-VL-72B, and Gemini-2.0-Flash—stay between roughly 55% and 61%, and several models score below chance on camera pose estimation. The paper argues that the two specific weak points are cross-view correspondence when objects are partially occluded and coarse camera pose estimation, and that chain-of-thought prompting does not reliably close the gap.","feed_headline":"MLLMs trail humans by 20+ points on multi-view scene questions","feed_subtitle":"New benchmark of 2,100+ multi-view questions puts best models near 60% against 82% human scores.","key_machinery":"The load-bearing object is All-Angles Bench itself: 90 curated real-world scenes with 2,132 three-option multiple-choice questions, each refined and verified by eight PhD annotators. Its central design device is paired-question generation—each of 85.3% of questions has a twin that rephrases the wording or swaps the referenced views while preserving the same visual correspondence. Comparing success on the two versions yields an 'inconsistent' (IC) score, which exposes whether a model's correct answer reflects robust multi-view understanding or a brittle inference. A second analytical instrument is a visualization prompt that asks the model to lay out objects and camera positions on a 10×10 grid; this reveals systematic misalignment in camera coordinates and viewpoint transformation.","core_discovery":"The central discovery, stated on the paper's own terms, is that current MLLMs remain far from human-level proficiency in multi-view understanding. On the 250-question human-evaluated subset, the human baseline is 82.0 average accuracy while the best model, InternVL2.5-38B, reaches 60.8, a gap of more than 20 points. The largest single failure is camera pose estimation: humans score 88.9, but the best MLLMs reach only the 30–50 range and many open-source models are at or below the 33.3% random-guess level. The paper also shows via a paired-question scheme that MLLMs often answer semantically equivalent rephrased questions inconsistently, meaning a correct answer is frequently a lucky guess rather than evidence of genuine multi-view understanding. The conclusion is that domain-specific refinements or modules that embed stronger multi-view awareness are necessary.","pith_inferences":["An immediate editorial test: recruit a larger panel of independent annotators (say 20 or more) for a fresh subset and measure inter-annotator agreement; if agreement is low on the hardest items, the 82.0 human baseline may be optimistic and the 20-point gap somewhat shrinks.","The paired-question design could be extended into a graded consistency metric—measuring how many view permutations and phrasing variants a model survives—which might predict reliability as an embodied agent better than raw accuracy.","A direct test of the paper's proposed remedy: fine-tune one of the evaluated open-source families on synthetic multi-view data with explicit 3D ground truth and see whether its All-Angles score jumps past the current 60.8 ceiling; this would separate the 'training data' explanation from the 'architecture' explanation."],"forward_implications":["All-Angles Bench gives the community a shared yardstick for multi-view reasoning, so claims about embodied or spatial competence in MLLMs can be checked against a fixed, human-verified question set.","Because camera pose estimation is where the biggest gap appears, camera-pose ordering can serve as a fast diagnostic test for multi-view geometric awareness in future model releases.","The paired-question inconsistency score shows that single-question accuracy overstates multi-view understanding; reporting both accuracy and consistency becomes the norm for such benchmarks.","The failure of chain-of-thought prompting to consistently help implies that the bottleneck is perceptual alignment across views rather than language-level reasoning, steering future work toward domain-specific training or geometric modules."],"supporting_citations":[{"why":"Supplies most of the 90 real-world multi-view scenes used as the substrate for question generation.","marker":"[18]"},{"why":"Supplies the remaining multi-view scenes with dense viewpoints, curated to 4-5 spatially dispersed views.","marker":"[24]"},{"why":"Serves both as the MLLM that generated draft questions and as a leading closed-source baseline that exhibits the human-model gap.","marker":"[40]"},{"why":"Defines the Gemini-2.0-Flash baseline, one of the strongest closed-source models, and is analyzed in the camera-pose visualization.","marker":"[45]"},{"why":"Provides the Claude-3.7-Sonnet baseline, which underperforms open-source models on orientation-sensitive tasks.","marker":"[1]"},{"why":"Defines Qwen2.5-VL-72B, an open-source model near the top of the benchmark, attributed to its video and grounding strengths.","marker":"[4]"},{"why":"Defines InternVL2.5-38B, the overall best-performing model in the paper, which anchors the upper bound of current MLLM performance.","marker":"[8]"},{"why":"Supplies the chain-of-thought prompting baselines and the 10x10 grid visualization prompt used to diagnose orientation-related failures.","marker":"[55]"}],"fun_headline_variants":["Multi-view MLLMs trail humans by 20+ points","MLLMs lag humans 20+ points on multi-view scene questions","Camera pose stumps MLLMs: humans 89%, best AI ~50%","All-Angles Bench: MLLMs far from human-level multi-view"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"For the measured 20-point gap to mean what the paper says, the human-written ground-truth answers must be correct on nearly every question, and the two-person human baseline of 82.0 must be a faithful estimate of human-level performance; if either fails, the gap and the conclusion could be inflated.","fun_headline_variants_meta":{"raw":{"variants":["Multi-view MLLMs trail humans by 20+ points","MLLMs lag humans 20+ points on multi-view scene questions","Camera pose stumps MLLMs: humans 89%, best AI ~50%","All-Angles Bench: MLLMs far from human-level multi-view"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001132,"raw_usage":{"total_tokens":4758,"prompt_tokens":1051,"completion_tokens":3707,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":667,"completion_tokens_details":{"reasoning_tokens":3626}},"tokens_in":667,"tokens_out":3707,"duration_ms":23545,"temperature":1.0,"reasoning_tokens":3626,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:28:03.866725+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a fresh panel of at least five independent annotators on a randomly selected 250-question subset, computing inter-annotator agreement; if their average accuracy falls more than 10 points short of the paper's 82.0 human baseline, or if any MLLM reaches within 5 points of the new baseline on a held-out set, the central gap claim would be overturned.","supporting_citations":[{"cited_title":"Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives","cited_arxiv_id":null,"evidence_quote":"Supplies most of the 90 real-world multi-view scenes used as the substrate for question generation."},{"cited_title":"Ego-humans: An ego- centric 3d multi-human benchmark","cited_arxiv_id":null,"evidence_quote":"Supplies the remaining multi-view scenes with dense viewpoints, curated to 4-5 spatially dispersed views."},{"cited_title":"gpt4o, 2024","cited_arxiv_id":null,"evidence_quote":"Serves both as the MLLM that generated draft questions and as a leading closed-source baseline that exhibits the human-model gap."}],"review_version":1}