{"id":"0e7db910-3232-4993-99d0-6ab97b31125f","arxiv_id":"2501.05031","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A benchmark and evaluation framework that tests LVLMs on embodied cognition from egocentric video, covering static, dynamic, and hallucination scenarios.","lead":"ECBench is a new benchmark for testing whether large vision-language models understand the world from a robot's first-person view, using 4,324 question-answer pairs across 30 embodied cognition dimensions. It adds robot self-awareness, dynamic scene perception, and hallucination tests that earlier benchmarks lacked, and it shows that today's best models still answer only about half the questions correctly.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The empirical rankings and the 'complete inability' conclusion for dynamic scenes rest on an unvalidated GPT-4o judge that also serves as filter constructor and top test subject; without human-judge agreement or bias checks, ECEval scores cannot support the headline claims.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: ECEval's GPT-4o judge is unvalidated against human judgment and is entangled with the benchmark's construction and the best-performing test model. This is the correct point to stress-test because every headline comparison, including the claim that models lack dynamic perception and the ranking of GPT-4o above Qwen2VL and LongVA, passes through the judge's scores. If the judge is biased or noisy, the quantitative conclusions are not reliable even though the dataset itself may remain a useful resource. The paper does contain independent support for the benchmark resource: human annotation, public data and code, and a reasonably detailed evaluation framework. Those are genuine contributions and should be credited. However, they do not resolve the validity of the scoring layer. The appropriate disposition remains conditional: the benchmark can be accepted as a resource, but the evaluation claims should be revised or qualified until the judge is validated. I therefore keep the reader's verdict unchanged.","tokens_in":19969,"tokens_out":4975,"duration_ms":55769,"concrete_test":"Take a stratified sample of roughly 200 model responses spanning model families (GPT-4o, Qwen2VL, LongVA, Video-LLaVA, human), question formats (open/closed), and score ranges. Have three independent human annotators score every sampled response with the exact ECEval rubric from Appendix 9, then compute human-majority versus GPT-4o judge agreement (Cohen's kappa or ICC), both overall and per model. Also rescore the same sample with a second strong judge (e.g., Claude or Qwen2VL) without model identity information and compare rankings. If GPT-4o-human agreement is below about 0.7, or if agreement is systematically higher for GPT-family answers, the reported rankings and the dynamic-scene inability claim require re-derivation with a validated or human-scored evaluation protocol.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claims—that mainstream LVLMs perform poorly in dynamic scenes, that robot-centric questions are harder, and that GPT-4o-[32f] is the best model—are all computed from ECEval scores produced by a GPT-4o judge (Appendix 9, 'Scoring'). The paper reports no human correlation, no inter-annotator agreement, and no bias check for this judge. This matters because GPT-4o is used in three roles: in Section 3.1 it filters questions via blind answering; in Section 4.5 and Appendix 9 it scores every model response; and in Table 2 it is also the top-scoring test model. If the judge is lenient toward GPT-4o-style concise answers, or toward the reference answer's phrasing, the numeric gaps and the 'complete inability' statement for dynamic scenes are not established. The problem is not fully mitigated by the 90.31% close-ended share, because binary semantic-equivalence judgments are still made by GPT-4o rather than by exact match. For open-ended questions, ECEval relies on manually annotated 0.5-point anchors, but no agreement statistics for these anchors are given either. Thus the entire model-comparison layer inherits an unvalidated scoring function, making the paper's evaluation conclusions load-bearing and currently unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"ECBench is a new embodied-cognition benchmark for evaluating LVLMs from egocentric RGB-D indoor videos, comprising 386 videos and 4,324 QA pairs across static-scene, dynamic-scene, and hallucination subsets organized into a 30-dimension taxonomy. The authors propose ECEval, a hybrid scoring scheme that uses GPT-4o for binary scoring of close-ended questions and multi-level scoring, anchored on manually annotated partial-credit answers, for open-ended questions. They report evaluations of twelve LVLM variants plus a blind LLM baseline and a human reference, and they conclude that mainstream LVLMs are weak in dynamic scenes and robot-centric questions, that embodied hallucination is widespread, and that GPT-4o-[32f] is the best model, scoring 50.35 overall versus 94.96 for humans.","tokens_in":20184,"tokens_out":10762,"duration_ms":98588,"significance":"If the evaluation layer is valid, ECBench fills a genuine gap: existing embodied QA benchmarks largely lack robot-centric, dynamic-scene, and hallucination dimensions. The construction is careful in several measurable respects: category-independent manual annotation, a blind GPT-4o filtering step to remove text-only-answerable questions, transparent and reproduced prompts for benchmarking and scoring, a human ceiling score, and public release of data and code. The blind-LLM baseline of 24.09 suggests the filtering did reduce common-sense leakage relative to OpenEQA's reported 33.5. However, every headline claim about model competence is only as trustworthy as ECEval, and ECEval's validity is currently not demonstrated; the empirical conclusions are therefore conditional on additional validation of the scoring function.","major_comments":[{"comment":"The central empirical claims, that models perform poorly in dynamic scenes, that robot-centric questions are harder, and that GPT-4o-[32f] is the best model, all depend on scores produced by GPT-4o under the ECEval protocol, yet no evidence is reported that the judge agrees with human judgments or is unbiased across model families. GPT-4o appears in three roles: blind filter in Sec. 3.1, judge in Sec. 4.5 and Appendix 9, and top-scoring test model in Table 2. The filter role is the least problematic, since the blind baseline is only 24.09 and the filter targets common-sense-answerable items; but the judge role is load-bearing. For open-ended answers, the 0.5-point anchors are manually annotated with no inter-annotator statistics reported; for close-ended answers, 'binary' scoring still requires a GPT-4o semantic-equivalence judgment rather than exact matching, so the 90.31% close-ended share does not remove the judge from the loop. The paper itself acknowledges in Sec. 4.5 that GPT-4o-based multi-level scoring is subject to GPT-4o biases. Without a human-correlation or agreement check, a systematic judge tendency, for example toward the concise phrasing encouraged by the Appendix 9 benchmarking prompt, could partly explain the reported 50.35 versus 24.09 gap and the model ordering in Table 2. The authors should report judge-human agreement (for instance Cohen's kappa on a sample), a bias analysis by model family and answer length, and ideally a second judge before the rankings are accepted.","section":"Sec. 4.5; Appendix 9 (Scoring); Sec. 3.1"},{"comment":"The conclusion of a 'complete inability' of LVLMs to perceive dynamic elements is drawn from a subset of only 248 QA pairs (Fig. 3c), split into four categories of roughly 60 items each; the 0.00 score of InternVL2-40B-[20f] on Quantity Dynamics and the low QD scores in general may be small-sample artifacts rather than evidence of a categorical inability. The per-category sample sizes are not reported in Table 2, and no confidence intervals are given. The authors should either enlarge the dynamic subset, which Appendix 10 acknowledges is currently limited, or reword the claims to refer to low performance on the available items, and zero-score observations should be accompanied by interval estimates.","section":"Sec. 4.3; Table 2; Fig. 3c"},{"comment":"No confidence intervals, standard errors, or significance tests are reported anywhere in the evaluation. Several load-bearing comparisons in Sec. 4.1 involve score differences of a few points, such as the 3.05-point gain attributed to increasing the frame count for Qwen2VL-72B (41.57 to 44.62), and the per-sub-ability cells in Tables 3-5 are even smaller. With QA-level variance and judge stochasticity, differences of this size may not be reliable; bootstrap confidence intervals over QA pairs, or a paired significance test, should be added before claims such as 'Reasoning problems impose higher demands on model capabilities compared to Perception questions' in Sec. 4.2 are made.","section":"Sec. 4.1; Tables 2-5"},{"comment":"The supplementary examples contain an internally inconsistent QA pair: in the second Missing Reference example under the ScanNet hallucination group, the question asks about a yellow bookshelf with stuffed toys in a third compartment, but the labeled answer states that there is no football under the desk or next to the sofa, which is the answer text of the preceding example. This looks like a copy-paste error, and if such mismatches exist in the released data, they contradict the 'stringent cross-validation' guarantee in Sec. 3.1 and could corrupt hallucination-subset scores. The authors should audit all QA pairs for answer-question mismatches and correct the example before publication.","section":"Appendix 6, Fig. 12"},{"comment":"The human reference score of 94.96 in Table 2 is central to the abstract's framing, but its measurement is not described anywhere: the number of participants, whether they were the same annotators who wrote the QA pairs, whether they saw the same frames and prompts as the models, and the agreement among raters are all unspecified. If the annotators who designed the questions also produced the human ceiling, the human score is likely inflated relative to an independent rater, which would exaggerate the model-human gap in the paper's headline comparison. A short protocol description and inter-rater agreement statistic should be added.","section":"Table 2; Sec. 4.1"}],"minor_comments":[{"comment":"The text states that Qwen2VL achieved 44.16, but no entry in Table 2 equals 44.16; the closest value is Qwen2VL-72B-[20f] at 44.62, so the cited number should be corrected or its provenance explained.","section":"Sec. 4.1 vs Table 2"},{"comment":"The filtering description says GPT-4o answers all questions six times and then 'this process is iterated thrice'; the intended protocol, three iterations of six repetitions or six runs per iteration, should be stated precisely.","section":"Sec. 3.1; Appendix 9"},{"comment":"The scoring prompt in Appendix 9 uses a 0-5 scale and defines anchors as the '5-score answer' and '3-score answer', while Sec. 3.1 describes a 0-1 scale in 0.2 increments; the mapping between the two conventions should be made explicit.","section":"Sec. 3.1; Appendix 9"},{"comment":"The axes and the combination of bar and overlaid distribution plots in Fig. 4 are not clearly labeled, and the units for question length and vocabulary size are ambiguous; axis titles and a legend should be added.","section":"Fig. 4"},{"comment":"The per-category sample sizes, especially the 248 dynamic-scene pairs and the small per-dimension counts in Fig. 3c, should be restated in the table captions or a footnote so that readers can assess the reliability of each cell.","section":"Table 2; Fig. 3c"},{"comment":"Reference [31] is formatted as 'R OpenAI' without author names; it should be replaced with the standard GPT-4 technical report citation.","section":"References"},{"comment":"The sentence 'All data and code is available' should read 'All data and code are available'.","section":"Abstract"},{"comment":"The comparison between the ECBench blind score (24.09) and the OpenEQA blind score (33.5) is presented as a 28% reduction in evaluation precision, but the two benchmarks differ in question-type mix, answer format, and difficulty; this comparison should be framed as suggestive rather than as a controlled measure of visual dependency.","section":"Sec. 4.1"}],"recommendation":"major_revision","confidential_remarks":"The benchmark resource and taxonomy are worthwhile contributions, and the construction pipeline is largely careful, but the evaluation methodology needs validation before the headline rankings can be trusted. The mismatched QA pair in Appendix 6 and the unexplained 44.16 in Sec. 4.1 suggest the released artifact and the final text should both be audited before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: ECBench is a solid, genuinely useful benchmark resource, but the paper's headline claims about model competence rest on an unvalidated GPT-4o judge that also served as the filter constructor and is the top-scoring test model. The benchmark itself deserves a serious look; the evaluation layer needs work before I'd trust the rankings.\n\nWhat's actually new: robot-centric self-awareness questions, a dynamic-scene subset capturing changes in and out of view, and embodied hallucination probes (counterintuitive scenes, unreliable user input) are not covered by OpenEQA or other existing benchmarks. The construction is careful: human annotation, category-independent balancing, a blind GPT-4o filter to remove text-answerable questions, and a combined binary/multi-level scorer with manually annotated 0.5-point anchors. The data and code are public. That's real value.\n\nSoft spots:\n\n1. The judge problem. GPT-4o is used to filter questions (Sec 3.1), to score every model response (Sec 4.5, Appendix 9), and is itself the best-performing model (Table 2). The paper reports no human correlation, no inter-annotator agreement, and no bias checks for the judge. With binary semantic-equivalence judgments made by the same model family being tested, the numeric gaps—and especially the \"complete inability\" claim about dynamic scenes—are not established. The 90.31% close-ended share doesn't fix this; the judge still decides correctness.\n\n2. Small dynamic subset. 248 QA pairs is thin for the strong conclusion that \"all mainstream LVLMs exhibit poor performance.\" Per-category dynamic scores rest on even fewer items; a few misjudged answers could move several points.\n\n3. Overstated comparisons. The blind LLM score comparison against OpenEQA (28% reduction) is confounded by question format: OpenEQA is entirely open-ended, ECBench is 90.31% close-ended. Lower blind scores on close-ended questions don't directly imply higher visual dependency.\n\n4. No error bars or significance tests anywhere. Some differences between models are large, but several (e.g., Qwen2VL-72B-8f vs 20f on some categories) are small enough that noise could matter.\n\nThe limitations section honestly admits missing proprietary models and limited dynamic-scene diversity, which is good. The core issue is that the evaluation scaffolding is not held to the same standard as the dataset construction.\n\nWho's this for: anyone building or evaluating LVLMs for embodied/egocentric video QA. The benchmark will likely get adopted because it fills a real gap. It deserves peer review—conditionally, with the judge validated against human ratings (or replaced by exact match on the close-ended portion) and claims softened accordingly.","headline":"Useful new embodied-cognition benchmark with genuinely new question types, but the evaluation layer needs a validated judge and toned-down claims before I'd trust the rankings.","tokens_in":20765,"tokens_out":2565,"would_cite":true,"duration_ms":24810,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ECBench is a 4,324-question benchmark showing that top vision-language models score only about half as well as humans on embodied cognition from egocentric video.","keywords":["embodied cognition","egocentric video understanding","vision-language models","dynamic scene perception","hallucination evaluation","robot-centric question answering","RGB-D video benchmark","multi-modal evaluation"],"falsifier":"Have independent human annotators re-score a random sample of 300 model responses with the ECEval rubric and compare their scores with GPT-4o's; if agreement (e.g., Cohen's kappa) is low, or if replacing GPT-4o with another judge changes the overall model ordering by more than a few points, then the benchmark's claim that models lack dynamic perception would not be established.","tokens_in":19734,"feed_emoji":"🤖","tokens_out":10254,"duration_ms":91009,"temperature":0.7,"pith_summary":"ECBench is built on the idea that before a vision-language model can reliably drive a robot, it must be tested on the embodied skills a robot actually needs—not just object recognition in static third-person scenes. The paper constructs a benchmark of 4,324 human-annotated question-answer pairs over 386 egocentric RGB-D videos, organized into 30 cognitive dimensions spanning static scenes, dynamic scenes, and hallucination. With it, the authors report that every model they test scores far below human level: the best, GPT-4o with 32 frames, reaches 50.35% overall while humans reach 94.96%, and a blind GPT-4o given no video still scores 24.09%. The central claim is that current models lack reliable first-person self-awareness, dynamic scene perception, and the ability to reject bad user instructions, and that a systematic benchmark is what exposes these gaps.","feed_headline":"Top vision-language model scores 50% on embodied cognition; humans 95%","feed_subtitle":"A 4,324-question benchmark shows models struggle most with dynamic scenes and robot self-awareness.","key_machinery":"The load-bearing machinery is the 30-dimensional capability taxonomy itself, which turns 'embodied cognition' into measurable question categories: 19 static-scene abilities split into scene-based and robot-centric cognition, four dynamic-scene categories (information, quantity, spatial, and state dynamics), and seven hallucination dimensions covering over-confidence in common sense and in user input. The taxonomy is enforced by class-independent human annotation—annotators are assigned quotas per ability rather than per video, which balances rare and common skills—and by a GPT-4o blind filtering loop in which the model answers every question without video six times, and annotators rework questions it can answer correctly, repeated three rounds, to keep the benchmark visually dependent. Scores come from ECEval, which uses binary 0/1 scoring for closed-ended questions and a 0-to-1 multilevel rubric anchored by a human-written 0.5-point reference answer for open-ended questions; this design is what lets the paper claim both precision and fairness in the final rankings.","core_discovery":"On the paper's own terms, ECBench is the first benchmark to break embodied cognition for large vision-language models into a fixed taxonomy of 30 abilities, and to evaluate them with mixed open and closed questions that cannot be answered from language priors alone. The empirical discovery is that capability is not uniform: robot-centric questions that require the model to reason about its own position, trajectory, or future actions are harder than scene-based third-person questions; dynamic scene questions, especially quantity dynamics, are nearly unsolved; and embodied hallucination splits into two failure modes—over-confidence in common sense and over-confidence in user input—with the latter barely addressed by any model. The strongest system, GPT-4o at 32 frames, still falls 44.61 points behind the human average, and the paper reads this as evidence that current LVLMs have third-person static-scene cognition but not yet first-person understanding in dynamic scenes.","pith_inferences":["Our inference: a fair next experiment is to fine-tune one open model on dynamic egocentric video and run ECBench before and after; if scores do not move, the benchmark may be testing something other than the visual dynamics it names.","Our inference: because ECEval uses GPT-4o as judge, rankings may be sensitive to judge choice; re-scoring a fixed set of responses with another capable model or with human raters would tell whether the reported ordering is robust.","Our inference: the benchmark's videos are under five minutes, so extending the same taxonomy to hours-long egocentric streams is a natural next step and would show whether self-localization and memory degrade further over longer horizons."],"forward_implications":["If ECBench is measuring what it says, then scaling input frames from 8 to 20 is not a sufficient cure for dynamic-scene failures; frame count helps some models, but the dynamic sub-scores stay near or below 25%.","The robot-centric results imply that first-person self-awareness (trajectory review, distance and azimuth awareness, movement imagery) is a distinct bottleneck: GPT-4o scores 49.04 there versus 59.74 on scene-based questions.","Hallucination results imply that current models cannot reliably correct user input: on missing, erroneous, and ambiguous references, most models score near zero, which matters for safe human-robot interaction.","The blind GPT-4o score of 24.09, compared with the 33.5 the paper reports for OpenEQA, implies ECBench is substantially less answerable from common sense alone, so score differences on it are more likely to reflect visual understanding."],"supporting_citations":[{"why":"OpenEQA defines the open-vocabulary embodied QA task that ECBench extends and supplies the 33.5 blind score used to argue ECBench has higher visual dependence.","marker":"[27]"},{"why":"MMBench-Video provides the systematic capability-taxonomy approach that ECBench adapts into 30 embodied cognition dimensions.","marker":"[9]"},{"why":"MVBench supplies the spatial-temporal task-construction methodology behind ECBench's dynamic scene categories.","marker":"[20]"},{"why":"GPT-4o system card documents the model used as blind filter, as ECEval judge, and as the best-performing evaluated system.","marker":"[13]"},{"why":"ScanNet supplies 140 real-world RGB-D scans that ground the static scene question set.","marker":"[7]"},{"why":"MultiScan supplies 51 additional real-world scans with articulated objects for scene-based questions.","marker":"[29]"},{"why":"HM3D provides the virtual indoor environments in which robotic agent videos were recorded.","marker":"[32]"},{"why":"Explore until confident provides the embodied question-answering agent whose videos contribute first-person robot footage.","marker":"[33]"},{"why":"OpenFMNav provides the object-navigation agent whose video streams make ECBench's robot behavior realistic.","marker":"[16]"}],"fun_headline_variants":["New embodied cognition benchmark: AI scores 50%, humans 95%","ECBench: AI lags humans by 44.6 points in embodied cognition","AI scores half on embodied cognition; humans 95% in new benchmark","Embodied AI test: GPT-4o hits 50%, humans 95% on ECBench"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline conclusion depends on the assumption that GPT-4o, used both to filter out questions answerable without video and to score every model answer, gives scores that match human judgment and do not favor any model family; the paper reports no human-agreement check for the judge.","fun_headline_variants_meta":{"raw":{"variants":["New embodied cognition benchmark: AI scores 50%, humans 95%","ECBench: AI lags humans by 44.6 points in embodied cognition","AI scores half on embodied cognition; humans 95% in new benchmark","Embodied AI test: GPT-4o hits 50%, humans 95% on ECBench"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001062,"raw_usage":{"total_tokens":4458,"prompt_tokens":955,"completion_tokens":3503,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":3416}},"tokens_in":571,"tokens_out":3503,"duration_ms":25460,"temperature":1.0,"reasoning_tokens":3416,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:19:42.854280+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have independent human annotators re-score a random sample of 300 model responses with the ECEval rubric and compare their scores with GPT-4o's; if agreement (e.g., Cohen's kappa) is low, or if replacing GPT-4o with another judge changes the overall model ordering by more than a few points, then the benchmark's claim that models lack dynamic perception would not be established.","supporting_citations":[{"cited_title":"Openeqa: Embodied question answering in the era of foundation models","cited_arxiv_id":null,"evidence_quote":"OpenEQA defines the open-vocabulary embodied QA task that ECBench extends and supplies the 33.5 blind score used to argue ECBench has higher visual dependence."},{"cited_title":"Mvbench: A comprehensive multi- modal video understanding benchmark","cited_arxiv_id":null,"evidence_quote":"MVBench supplies the spatial-temporal task-construction methodology behind ECBench's dynamic scene categories."},{"cited_title":"Chang, Manolis Savva, Maciej Hal- ber, Thomas A","cited_arxiv_id":null,"evidence_quote":"ScanNet supplies 140 real-world RGB-D scans that ground the static scene question set."},{"cited_title":"Chang, and Manolis Savva","cited_arxiv_id":null,"evidence_quote":"MultiScan supplies 51 additional real-world scans with articulated objects for scene-based questions."},{"cited_title":"Turner, Eric Undersander, Wojciech Galuba, Andrew West- bury, Angel X","cited_arxiv_id":null,"evidence_quote":"HM3D provides the virtual indoor environments in which robotic agent videos were recorded."},{"cited_title":"Openfmnav: Towards open-set zero-shot object navigation via vision- language foundation models","cited_arxiv_id":null,"evidence_quote":"OpenFMNav provides the object-navigation agent whose video streams make ECBench's robot behavior realistic."}],"review_version":1}