{"id":"e3ebb9a5-05e5-4301-8f25-129974b80183","arxiv_id":"2505.09990","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A three-stage benchmark consisting of 982 pointing tasks, a live pairwise arena with 4,500 votes, and a robot manipulation study shows that pointing-supervised open models such as Molmo-72B can match proprietary models, with high reported correlations between stages.","lead":"PointArena is a new benchmark for testing whether AI models can point to the correct part of an image in response to a language instruction. It combines a curated dataset, a live human-vote arena, and a real robot, and reports that models trained on pointing data outperform or match proprietary models.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The R²=0.92 proxy claim rests on only three agents including a human reference; with n=3 it cannot establish that Point-Bench predicts real-world pointing success.","rationale":"The paper's main contribution is an evaluation platform, and its strongest claim is that static benchmark accuracy transfers to real-world task success. Section 4.3 uses R²=0.92 from exactly three agents, including a human reference. With n=3, a linear model can fit the data closely regardless of the true relationship, and the human reference is qualitatively different from the MLLMs being benchmarked. This is not a disagreement with consensus but a matter of insufficient independent evidence: the number of units is too small to warrant the predictive claim. The Section 3.2 acceptance rule compounds this by filtering benchmark queries through current model failures, making Point-Bench a deliberately hard, non-random sample; mapping it to real-world success therefore requires stronger external validation than three points. The Molmo-72B 'highest' result is also a 0.43 pp gap with p≈0.29, so the headline ranking is fragile, though the benchmark-validation issue is more load-bearing. The CONDITIONAL verdict remains appropriate: the platform is a reasonable artifact, but the strong correlation and ranking claims need more data. A concrete test is to run Point-Act with several more models on the same scene and refit without the human reference; if the correlation collapses, the proxy claim should be weakened.","tokens_in":12422,"tokens_out":4880,"duration_ms":52905,"concrete_test":"Run Point-Act with at least six additional MLLMs spanning the Point-Bench accuracy range on the same fixed scene, then refit the Point-Bench-to-robot-success regression excluding the human reference and report the bootstrap 95% confidence interval for R² and the Spearman rank correlation. If the rank correlation is not significant or R² drops below 0.7, the current R²=0.92 is a three-point artifact and the predictive-validity claim should be softened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.3 claims 'Point-Bench accuracy predicts real-world task success' with R²=0.92, based on Point-Act data from exactly three agents: Molmo-7B-D, GPT-4o, and a human reference, on one fixed scene with 10 participants and 3 trials each. With only three points, the linear regression has one degree of freedom after fitting, so a high R² is almost guaranteed and provides no information about predictive validity. The human reference anchors the high end of both scales and likely drives the fit; excluding it leaves two model points, for which the correlation is undefined. Thus the central claim that Point-Bench is a reliable proxy for real-world pointing is presently unsupported. The Section 3.2 acceptance rule (keep queries only if one or fewer of three current MLLMs answer correctly) also makes dataset composition depend on the filter models' failure modes, so the static benchmark's representativeness is not independently established by this three-point correlation. The Molmo-72B 'highest' claim is likewise a 0.43 percentage-point gap with p≈0.29, so the headline ranking is fragile, but the benchmark-validation issue is more load-bearing because the other conclusions inherit it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces PointArena, a three-part evaluation platform for language-guided pointing in multimodal models: Point-Bench, a curated 982-item static dataset spanning five reasoning categories; Point-Battle, a live pairwise human-vote arena; and Point-Act, a real-world robotic manipulation setup with a user study. The authors benchmark 16 MLLMs, report that Molmo-72B is the top Point-Bench performer, claim that explicit pointing supervision boosts accuracy, and report strong correlations among the three evaluation stages, including an R^2=0.92 claim that Point-Bench accuracy predicts real-world task success.","tokens_in":12638,"tokens_out":6416,"duration_ms":59924,"significance":"If the results hold, PointArena would be a useful community resource: the dataset is curated and publicly available, the three-stage design is thoughtful, the zero-shot evaluation protocol is standard, and the authors report standard deviations over three runs. The inclusion of a live arena and a real-robot evaluation is a strength, and the paper honestly discloses the statistically insignificant Molmo-Gemini margin. However, the load-bearing validation of Point-Bench as a proxy for real-world pointing rests on a three-point regression, and several headline claims are stronger than the evidence supports, so the current version does not fully substantiate the central conclusions.","major_comments":[{"comment":"The claim 'Point-Bench accuracy predicts real-world task success' with R^2=0.92 is based on exactly three agents (Molmo-7B-D, GPT-4o, and a human reference) evaluated on one fixed scene. With n=3, a linear regression has one residual degree of freedom, so a high R^2 provides essentially no information about predictive validity. The human reference anchors the high end of both scales and is likely responsible for much of the fit; excluding it leaves only two model points. This does not support the load-bearing conclusion that Point-Bench is a reliable proxy for real-world pointing. Please collect data from additional agents and scenes, or explicitly reframe this as an illustrative pilot rather than a validation.","section":"§4.3 (Point-Act validation)"},{"comment":"The text states that Molmo-72B outperforms Gemini-2.5-Pro by 0.43 percentage points and calls this margin 'statistically insignificant (p≈0.29)', yet the Abstract says Molmo-72B 'consistently outperforms' other models and §4.2 says it 'achieves the highest performance on the Point-Bench benchmark'. These statements are internally inconsistent. The headline ranking should be reported as statistically tied at the top, with appropriate multiple-comparison correction if a ranking is still claimed.","section":"§4.2 and Abstract (Molmo-72B ranking)"},{"comment":"The claim that 'explicit pointing data is a key driver of model accuracy' is supported by comparing Qwen2-VL-7B (17.4%) with Qwen2.5-VL-7B (52.3%). These are different model generations that differ in architecture, pretraining data, and other training choices, not solely in the presence of PixMo pointing data. This confound does not justify the causal attribution of the large gain to pointing supervision. Provide a controlled comparison using the same base architecture with and without the pointing corpus, or substantially weaken the claim.","section":"§4.2 (Pointing supervision)"},{"comment":"The Point-Bench construction accepts a query only if one or fewer of three anonymized MLLMs answer correctly. This makes the dataset composition depend on the failure modes of the filter models, whose identities are not disclosed. Since several evaluated models come from the same families as the likely filter models, the scoring on the resulting subset may be biased for or against particular models. This compromises the neutrality of the benchmark and the generality of conclusions about category difficulty. Please disclose the filter models and analyze the sensitivity of the results to the acceptance threshold and to the choice of filters.","section":"§3.2 (Dataset acceptance rule)"}],"minor_comments":[{"comment":"'collected from public sources posted after 20 April 20, 2025' appears to be a typo; the intended date is probably 'April 20, 2025'.","section":"§3.2"},{"comment":"The text refers to 'GPT-o3 improved by 21.1 points over GPT-4-Turbo' but Figure 5a labels this as GPT-4.1; the text also mentions 'Gemini-2.5-Flash' while the figure caption says 'Gemini-2.0-Flash'. Please align model names between text and figures.","section":"§4.2 and Figure 5a"},{"comment":"The Deitke et al. 2024a and 2024b entries appear to be the same paper with slightly different subtitles; please merge or correct the duplicate.","section":"References"},{"comment":"The phrase 'double-blind MLLM' is ambiguous; it likely means that the user (not the model) is blind to the agent identity. Please reword for clarity.","section":"§3.4"},{"comment":"The R^2=0.85 regression is computed from only five models; please include confidence intervals and explicitly acknowledge the small sample size in the text.","section":"Figure 5b"},{"comment":"The description of Point-Bench as 'the largest benchmark for evaluating language-guided pointing' is asserted without a comparison to other benchmarks; please cite or qualify this claim.","section":"§1 and §3.2"},{"comment":"There are inconsistent model-name spellings such as 'LLaV A' (LLaVA) and 'Molmo-7B-O'; please standardize the naming.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper comes from a prominent group and the platform is a useful community contribution, but the central validation claims need substantive revision: the n=3 regression cannot support the proxy claim, the Molmo-72B 'consistently outperforms' statement conflicts with the authors' own p-values, and the supervision-boost claim is confounded across model generations. These issues are fixable with additional experiments or appropriately weakened claims, so major_revision is appropriate rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, PointArena is a genuinely useful benchmark artifact: 982 curated pointing tasks across five categories, a live blind arena with over 4,500 votes, and a robot execution stage. Second, the headline validation claim—that Point-Bench accuracy predicts real-world task success (R²=0.92)—is not supported by the evidence; it is a linear fit to three agents, one of which is the human reference. That overclaim should not sink the paper, but it needs to be fixed.\n\nWhat's new: most grounding benchmarks like RefCOCO stop at object localization; PointArena broadens to affordance, counting, steerable, and reasoning categories. The three-stage pipeline (static, human preference, robot) is a reasonable way to triangulate pointing quality. The evaluation covers 16 models with three runs each. I buy the directional findings: pointing-supervised models do better (Qwen2.5-VL vs Qwen2-VL is a large jump), open models like Molmo compete with proprietary ones, and CoT hurts pointing. The authors also honestly list limitations about SAM boundaries and contamination.\n\nSoft spots, in order of severity. 1) The R²=0.92 correlation is fitted to three points—Molmo-7B-D, GPT-4o, and a human oracle. With n=3, any monotonic pattern will look great; the human anchor does most of the work. Calling Point-Bench 'a reliable proxy' for real-world performance is a strong claim from a three-agent, one-scene study. 2) The Molmo-72B 'consistently outperforms' line in the abstract is not supported: the margin over Gemini-2.5-Pro is 0.43 percentage points with p≈0.29. The paper elsewhere admits the margin is statistically insignificant, so the abstract overstates. 3) The Section 3.2 acceptance rule keeps only queries that one or fewer of three current MLLMs answer correctly, so dataset composition depends on those models' failure modes. That is typical of hard benchmarks, but it means the benchmark is not a neutral sample of pointing tasks. 4) No data or code is linked in the preprint, which limits independent checks of the masks and arena. The citation pattern is fine—standard references, no obvious omissions.\n\nWho this is for: people building or evaluating multimodal models for pointing, especially robotics and assistive AI. The benchmark itself is worth having. I'd cite it for the dataset and the five-category decomposition, but not for the R² claim.\n\nRecommendation: this deserves peer review, not desk rejection. A referee should push the authors to either add more agents and scenes to the Point-Act validation or soften the proxy claim, and to fix the abstract's ranking language. With those changes it is a solid benchmark paper.","headline":"Solid new pointing benchmark with overclaimed validation; R²=0.92 rests on three agents and the top-model edge is within noise.","tokens_in":13202,"tokens_out":2337,"would_cite":true,"duration_ms":21978,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces PointArena, a three-stage platform of static benchmark, human-preference arena, and real-robot manipulation for evaluating multimodal models' language-guided pointing, and reports that Molmo-72B leads while…","keywords":["multimodal large language models","language-guided pointing","visual grounding benchmark","human preference evaluation","robotic manipulation","spatial reasoning","referential localization","pointing supervision"],"falsifier":"Run a broader Point-Act validation with at least ten diverse multimodal models and several tabletop scenes, then compare model ordering by Point-Bench score with ordering by physical pick-and-place success rate; if the ranking correlation drops well below the reported $R^2=0.92$, the claim that Point-Bench predicts real-world pointing would be falsified.","tokens_in":12230,"feed_emoji":"👉","tokens_out":8985,"duration_ms":85064,"temperature":0.7,"pith_summary":"PointArena's central claim is that language-guided pointing can be measured as a standalone capability, and that a static pointing benchmark can stand in for both human preference and real physical task success. The paper builds three linked evaluation stages: Point-Bench, 982 curated image-query pairs spanning Spatial, Affordance, Counting, Steerable, and Reasoning; Point-Battle, a blind pairwise human-voting arena with more than 4,500 votes; and Point-Act, a real xArm robot that executes pick-and-place from a model's predicted points. Testing 16 multimodal models, it reports that Molmo-72B scores highest on Point-Bench, that explicit pointing supervision lifts accuracy far more than model scale does, and that Point-Bench accuracy correlates with human preference at $R^2=0.85$ and with robot task success at $R^2=0.92$. If these correlations hold, a cheap static test can rank models for embodied and assistive use, and pointing should be treated as its own training target rather than an emergent side effect of general vision-language training.","feed_headline":"Molmo-72B tops new pointing benchmark with robot validation","feed_subtitle":"Static pointing scores and real-world robot success align closely across PointArena's three evaluation stages.","key_machinery":"The machinery is a single success rule applied across all stages: a predicted point is correct if it falls inside a human-verified binary target mask, and a task succeeds if the predicted set covers every target region. That rule makes Point-Bench automatically scoreable, makes model outputs directly comparable in Point-Battle, and converts naturally into robot actions in Point-Act because each point is just an image coordinate. The benchmark's five categories (Spatial, Affordance, Counting, Steerable, Reasoning) are the sampling mechanism for covering different grounding demands, and a query is admitted to the dataset only if one or fewer current multimodal models answer it correctly, which keeps the benchmark hard while making its composition depend on the failure modes of the models it evaluates.","core_discovery":"The paper's discovery is that pointing is a separable, trainable capability with external validity. On the 982-example Point-Bench, Molmo-72B outperforms all other tested models, with Gemini-2.5-Pro statistically tied; open-source models trained on explicit pointing data (Molmo, Qwen2.5-VL with PixMo) match or beat proprietary models, while LLaVA variants without such data land at 4.8–17.4%. The same predicted points drive all three stages, so the authors can compare static accuracy, human preference, and physical execution directly: Point-Bench agrees with Point-Battle at $R^2=0.85$ and with Point-Act at $R^2=0.92$, and Molmo-7B-D beats GPT-4o on the robot by 65% in their user study. The ablation results add a second finding: chain-of-thought prompting hurts pointing accuracy in both GPT-4o and Gemini-2.5-Flash, suggesting that spatial pointing benefits from tight, coordinate-oriented prompts rather than extended verbal reasoning.","pith_inferences":["Editorial inference: If the $R^2=0.92$ proxy relationship holds beyond the three evaluated agents, Point-Bench could serve as a low-cost screening gate for robot manipulation policies, reserving expensive physical trials for models that already pass the static threshold.","Editorial inference: The chain-of-thought finding suggests spatial pointing is more perceptual than deliberative; a testable extension is whether visual backbones pretrained on dense alignment tasks benefit more from additional pointing data than from reasoning prompts.","Editorial inference: Because dataset acceptance depends on current model failures, Point-Bench is inherently a moving target; future releases may need a rolling refresh from Point-Battle's user-uploaded images to avoid contamination and saturation.","Editorial inference: The benchmark's category structure could be reused to build curriculum data for pointing supervision, with per-category scores identifying which spatial skill a model lacks."],"forward_implications":["Model releases can be screened cheaply: Point-Bench accuracy on 982 pairs gives a first-pass estimate of how well a model will point in physical pick-and-place settings.","Explicit pointing supervision is a stronger lever than parameter count; scaling from 7B to 72B changes accuracy by only a few points in the tested open-source families.","Prompt engineering for pointing should stay concise: adding chain-of-thought or verbose user-style phrasing hurts spatial grounding in GPT-4o and Gemini-2.5-Flash.","Arena-style human-preference rankings can track progress as static benchmarks saturate, because the two evaluation modes align at $R^2=0.85$.","Open-weight pointing models such as Molmo-7B-D can outscore proprietary APIs in human preference and robot usability, making them viable for assistive and embodied applications."],"supporting_citations":[{"why":"Supplies the PixMo corpus and Molmo models used to test whether explicit pointing supervision boosts accuracy.","marker":"Deitke et al. [2024a]"},{"why":"Provides the Molmo architecture and pointing coordinate predictions that achieve top Point-Bench and Point-Battle results.","marker":"Deitke et al. [2024b]"},{"why":"RoboPoint spatial-affordance data is cited as the spur for post-December-2024 model gains and is the robotic pointing baseline Point-Act builds on.","marker":"Yuan et al. [2024]"},{"why":"Chatbot Arena's Elo methodology is adopted for Point-Battle's blind pairwise voting.","marker":"Chiang et al. [2024]"},{"why":"Segment Anything generates the initial masks that annotators refine into Point-Bench ground-truth targets.","marker":"Kirillov et al. [2023]"},{"why":"RefCOCO defines the referring-expression object-localization task that Point-Bench extends beyond.","marker":"Kazemzadeh et al. [2014]"},{"why":"Chain-of-Thought prompting is the ablation condition that the paper shows harms pointing accuracy.","marker":"Wei et al. [2022]"}],"fun_headline_variants":["Pointing benchmark links static scores to real robot success","Molmo-72B wins pointing test; robot users beat GPT-4o","Explicit pointing training boosts models; long reasoning hurts","PointArena: humans judge pointing, robots confirm it works"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that Point-Bench accuracy predicts real-world pointing ability, with the evidence being a linear fit to only three agents—Molmo-7B-D, GPT-4o, and a human reference—tested on one fixed scene with ten participants.","fun_headline_variants_meta":{"raw":{"variants":["Pointing benchmark links static scores to real robot success","Molmo-72B wins pointing test; robot users beat GPT-4o","Explicit pointing training boosts models; long reasoning hurts","PointArena: humans judge pointing, robots confirm it works"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000154,"raw_usage":{"total_tokens":1240,"prompt_tokens":1004,"completion_tokens":236,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":620,"completion_tokens_details":{"reasoning_tokens":165}},"tokens_in":620,"tokens_out":236,"duration_ms":2762,"temperature":1.0,"reasoning_tokens":165,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:18:55.712578+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a broader Point-Act validation with at least ten diverse multimodal models and several tabletop scenes, then compare model ordering by Point-Bench score with ordering by physical pick-and-place success rate; if the ranking correlation drops well below the reported $R^2=0.92$, the claim that Point-Bench predicts real-world pointing would be falsified.","supporting_citations":[],"review_version":1}