Pith. sign in

Paper Citation Record · LEDGER

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

As of 23 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 16 inbound Pith citation observations for arXiv:2505.09990.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.09990 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:24:01.169260Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:20:26.059257Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:29:35.800189Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2219b9ac-df78-42ae-9265-3b8a45fb70c8 · outbound

This paper cites GPT-4 Technical Report.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:00.983152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:00.983152Z digest=sha256:d2bb22c43ffb97a8a055466793b6d33e3dbe6faee1bc8e496928c8b5300d8cc5

Observation dee1d29d-305f-471b-b72e-f55f77c9c521 · outbound

This paper cites an unresolved cited work.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:24:01.752211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T21:24:00.988229Z digest=sha256:181298fa39052fdd21fbf9d6e6bc28c1f215dd9b17098dcf16593ddbb1674884

Observation 9ad60340-6412-4c9f-b267-5d6fb77174b0 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing The claude 3 model family: Opus, sonnet, haiku

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:24:01.739207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T21:24:00.992592Z digest=sha256:f8677d2a698c6273cc7f145e2b67b87ab224e055e80d328055d3db8d006e310b

Observation 6dc0f7fa-e12b-442d-ad87-56222dca88da · outbound

This paper cites Qwen2.5-VL Technical Report.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:00.997319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:00.997319Z digest=sha256:0ef9408e5b3fdc52ae8517485b983962090ffcba8bfeea40caca79563542c3c4

Observation cda1af7f-3927-41c0-88ad-921203d727a5 · outbound

This paper cites Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.002029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.002029Z digest=sha256:24f034b18b67b07afa3a1bd56fafac22cf72c8afb55537255dc9d6e9b8ba9f3c

Observation 81b51eda-55be-4b02-a2fa-93d594d6343e · outbound

This paper cites Vizwiz: nearly real-time answers to visual questions.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Vizwiz: nearly real-time answers to visual questions

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:24:01.725088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T21:24:01.006649Z digest=sha256:dfea1e89d50062415ee4afe744e5463634e17f4c3ccd6e0615a00f13b4a8fa88

Observation 8449394f-2901-494b-b8e9-1541c3e67514 · outbound

This paper cites ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.011599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.011599Z digest=sha256:5f951903ef11cdc9463b22c18cfed62e7267682dc417290534082b2ee65402b5

Observation 632fb135-1918-402e-a31a-5a11440682d5 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.016035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.016035Z digest=sha256:a6895ed487c2173f9828a7e3ffbb69791371d6a3519ec9b735a53fe2672cbf4f

Observation 9be612a2-71b3-4ef0-b76f-7e313b9d82ff · outbound

This paper cites GuessWhat?! Visual object discovery through multi-modal dialogue.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing GuessWhat?! Visual object discovery through multi-modal dialogue

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.020494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.020494Z digest=sha256:37434608fdacba194611d5eaf9898be8559df5cff955a0e75b22e4b37c71099d

Observation cf4f6c35-192a-4d9f-8515-eea40efbd1dd · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.029358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.029358Z digest=sha256:b104ee65eb565b3922c1460604e0f67e6f2701fa96b453c7903a0a6dc6eea9a4

Observation e3862487-399c-48ed-a95d-aeece23da61e · outbound

This paper cites AR2-D2:Training a Robot Without a Robot.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing AR2-D2:Training a Robot Without a Robot

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.034080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.034080Z digest=sha256:a8c008107ca73ba73b3ea47a65337e5099f23d0397e1e3acda2d9caca93f0c9b

Observation 63440c4a-c670-41f1-904c-b6789f1fb571 · outbound

This paper cites Manipulate-Anything: Automating Real-World Robots using Vision-Language Models.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Manipulate-Anything: Automating Real-World Robots using Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.038681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.038681Z digest=sha256:0b42b30ec313d525c1886e6c29622406652780d68d163462782d8c5dbe5b3448

Observation f813e39e-b487-4ea8-95d2-fbbf9ec92f9f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.043011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.043011Z digest=sha256:7f1276d9b1682776d59dfdd36a1500adc68a5adebd3f1aa56dd448837d263d80

Observation d706436b-bbe4-47d2-a168-f184c2000066 · outbound

This paper cites A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.047565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.047565Z digest=sha256:4f160a1d5e169b31717e30e2d42801c440c3669ca2786d82dff7a65b094bff65

Observation c913113c-306a-4268-b462-da67fe257c52 · outbound

This paper cites Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.052167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.052167Z digest=sha256:67a15280b2f769b0fa6c5c474852ec51a713470325c00618537dc49c63871180

Observation 7d9ca603-53b9-4985-b0f9-785d1f42c736 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Referitgame: Referring to objects in photographs of natural scenes

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.057019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.057019Z digest=sha256:b2fd2b949c20d563b993ede9bc70d5e61e21d0b3d91802e6a8f804c628aaee48

Observation 0ebfc7ee-b121-4fa5-9a4a-7f1e487e0763 · outbound

This paper cites Tournament evaluation of large language models, 2025.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Tournament evaluation of large language models, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:24:01.702253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T21:24:01.062034Z digest=sha256:89f0a446ce26dc5c47f955aa2e528f0879a58f5d75cc88e7ec459c63b56238ae

Observation 11a55da1-0788-4a55-8ef5-9665e2b508a7 · outbound

This paper cites Segment Anything.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Segment Anything

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.066586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.066586Z digest=sha256:2dd45b427bfed67bb0e81f2b7243e8377566776ba96a207c95f45bf2f3a98e9c

Observation 54b65a26-b2a7-4707-8519-e3abb3f8ca96 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing LLaVA-OneVision: Easy Visual Task Transfer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.071722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.071722Z digest=sha256:32eec22df0924828f65a71b72f51b2fa4d58a4597893661850cc45882aed4986

Observation 12678563-b8c7-4c75-8dd0-f305f3bdb733 · outbound

This paper cites am-ELO: A Stable Framework for Arena-based LLM Evaluation.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing am-ELO: A Stable Framework for Arena-based LLM Evaluation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.076251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.076251Z digest=sha256:ef1ebac4e044afeb77d3280c5dcd665f48bcb12ea7ddd586acad78fc7d0e777c

Observation b74a53f2-a5f5-4ddc-b1b6-c0683c186ad7 · outbound

This paper cites Improving Your Model Ranking on Chatbot Arena by Vote Rigging.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Improving Your Model Ranking on Chatbot Arena by Vote Rigging

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.080799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.080799Z digest=sha256:94b06166eeac783d80bb205fe406706cbadca32609148375c7ed089c4fc01091

Observation 7869ea04-7d41-4517-8449-17f882bea05e · outbound

This paper cites CityRefer: Geography-aware 3D Visual Grounding Dataset on City-scale Point Cloud Data.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing CityRefer: Geography-aware 3D Visual Grounding Dataset on City-scale Point Cloud Data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.085663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.085663Z digest=sha256:56ce4e1b2696f91bc0d5217abc5fac82f03335b869977e14ce87c6ce3af611e0

Observation ec0e5f45-3049-49ce-b7bd-a866edb56ba4 · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.090778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.090778Z digest=sha256:fc64f5d9d44d73d3d88d6140a120b3bc867c02b801952699d37c10d0c0d9fe7d

Observation a0a08dfc-a950-42ab-99cc-707b47749c09 · outbound

This paper cites Multimodal Explanations: Justifying Decisions and Pointing to the Evidence.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Multimodal Explanations: Justifying Decisions and Pointing to the Evidence

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.095790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.095790Z digest=sha256:b6e461bfc2587e471d24ff27d465f37bda40ff43f667296b42ff28b1fed4b207

Observation 878890d2-951f-4555-9dd9-6a3e351a436f · outbound

This paper cites Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.101202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.101202Z digest=sha256:047092fd5469a7a52830b69783756602fddd350d002266ffa28b3340e425b21a

Observation a35e52b8-ba94-426b-ad5e-8fbaf3dc8032 · outbound

This paper cites Sat: Spatial aptitude training for multimodal language models.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Sat: Spatial aptitude training for multimodal language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.106101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.106101Z digest=sha256:ff243281500f852368bdf4dfdfa8ee3cb5f947b45e4ef44b33dccaba1ba4c32d

Observation c4cb9e7e-9e74-40e8-8ec7-66cd1e6ccd70 · outbound

This paper cites Openarena: An open platform for llm-as-a-judge evaluation.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Openarena: An open platform for llm-as-a-judge evaluation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:24:01.688430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T21:24:01.111185Z digest=sha256:6b384874f05ab716b8e1d39d76ac734fe2d68c0b5f01b4aeafef5bdf941bfbda

Observation 66b74e56-b0f8-401c-a958-bd3afd85dd1c · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Gemini Robotics: Bringing AI into the Physical World

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.115657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.115657Z digest=sha256:3b1bb79af9030881ca09dec012c04101fbc23f057554f7289b581c7bd6ca21cb

Observation 38c0aae0-0a95-45c3-9736-81b688d373e1 · outbound

This paper cites A new look at infant pointing.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing A new look at infant pointing

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:24:01.675192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T21:24:01.121912Z digest=sha256:657c878d39acfbc0eca165e289fe6bd1d47e86e12c034b0b3e73dfa00805b3cd

Observation b1e0e513-6bbe-46e9-8132-81fb3f4c6289 · outbound

This paper cites A Comprehensive Review of Multimodal Large Language Models: Performance and Challenges Across Different Tasks.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing A Comprehensive Review of Multimodal Large Language Models: Performance and Challenges Across Different Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.126616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.126616Z digest=sha256:7eda0a9daecf27a980e02e6d73ad6bb6de5dca9894ab926ec6a2c929d147dd27

Observation afaaef75-58a1-4208-b2f1-c4ca25ba31a8 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Chain-of-thought prompting elicits reasoning in large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.131189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.131189Z digest=sha256:a9f975366bc440882d56193b688efc6ed8c626d20310ee50bca96473e5402f6f

Observation dd9ccb08-3e39-43a0-9298-1f57216104a4 · outbound

This paper cites Grok-2 model card.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Grok-2 model card

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:24:01.653072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T21:24:01.135323Z digest=sha256:c304409f4f13845c3ca261e8973fe68807b1f1f65414b18ac4ea252fbf733332

Observation 0a5b8dcc-b802-46d4-a45d-f2ab9418877f · outbound

This paper cites A survey on multimodal large language models.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing A survey on multimodal large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.139725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.139725Z digest=sha256:769a24aea1ecdeff3702b31e9db0c3a7edc2834b520729a4467e36509c4b05a9

Observation 0f848a5d-4384-4d58-80ae-137b1b0428e9 · outbound

This paper cites Modeling Context in Referring Expressions.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Modeling Context in Referring Expressions

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.149740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.149740Z digest=sha256:b4b8cccdbf73ff1b3d9a8f35103b87a53fc18e6ade9da7f4f66ea186fc1aabbe

Observation b1630d39-e1d2-42ae-9e1e-1d2b95845a11 · outbound

This paper cites RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.154135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.154135Z digest=sha256:9e9fcc4adb9781cd93a365f6b29f18285c4c7765b174ffa4a0f42e402c1641e6

Observation d42ccffb-65a6-4b6b-b42a-0822ff0a10db · outbound

This paper cites Robopoint: A vision-language model for spatial affordance prediction in robotics.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Robopoint: A vision-language model for spatial affordance prediction in robotics

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:24:01.639467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T21:24:01.159064Z digest=sha256:6c2325f704db6387af739508f031d27886700a8624a51cbee5d48971aa4ee8ad

Observation c6da27ed-7ce2-45c4-888c-e085982500bb · outbound

This paper cites Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.164150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.164150Z digest=sha256:4d69fa6af65b065eb66e203a4eb5907f62b111fc625660a67940a2075e56895f

Observation fe170848-47af-48fe-8b08-2a2d80520cb2 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.169260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.169260Z digest=sha256:6b755cacfc311729a3e45a494240afabcb4ae3c06c7bfa501e331ef95cd4dae7

Pith citing papers

Observation 73db2bcf-1a51-4c74-a4f5-8d545ab2785a · inbound

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding cites this paper.

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:21:29.777883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T04:21:29.526008Z digest=sha256:216f3d65d719a371c9a38947fdae7a13def3a3dca72c2db85e5c12c5a1baeb71

Observation baddc7ca-87d9-45e0-86da-474f4af62f9c · inbound

Seed1.8 Model Card: Towards Generalized Real-World Agency cites this paper.

Seed1.8 Model Card: Towards Generalized Real-World Agency PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T07:45:14.321140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T07:44:02.827006Z digest=sha256:15914b64f68568d679365bceed3a025db94827f0b5fe6822da1faf2e6397a344

Observation 3ce27e9b-59bd-4986-a77a-a854a49b07d5 · inbound

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents cites this paper.

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:26.216036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T13:49:49.730666Z digest=sha256:fea9afea5d5b1b477e02aacbe4156565a986f123d018da4afb2582053dc8c22a

Observation b7872288-c099-44fe-af46-6b1937304fba · inbound

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents cites this paper.

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:06:22.836095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T03:26:08.650099Z digest=sha256:4ad579e350145c8d04b5ea038737cbffd49492a3b9575b309f37e586df25e809

Observation ccd375ba-116a-4bfe-b769-a790e05b02b3 · inbound

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents cites this paper.

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:32:30.501953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T07:28:45.811192Z digest=sha256:3eabb6227153f66dca5881497e6ad5f7e130d9a43cf1636b6f16f939c50cd7c7

Observation 8d918e7f-bd15-4dbe-bdea-45df91653455 · inbound

MolmoAct2: Action Reasoning Models for Real-world Deployment cites this paper.

MolmoAct2: Action Reasoning Models for Real-world Deployment PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:06:05.132338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T17:53:44.901684Z digest=sha256:f125f5cba3d64c8a6e38e880e142c0d622c429f9c1fa38457abef54ab16080ba

Observation 9fc55664-f5a3-49a2-9a06-8e511421e035 · inbound

MolmoAct2: Action Reasoning Models for Real-world Deployment cites this paper.

MolmoAct2: Action Reasoning Models for Real-world Deployment PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.197919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:0af70bfb7b0e1cf7a69424ac9762d731d479e0bed6d4a8845e33febbe4bc39b2

Observation 85f87663-0c8f-42a6-b2dd-83a67e1e51be · inbound

ZAYA1-VL-8B Technical Report cites this paper.

ZAYA1-VL-8B Technical Report PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 146

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:21:23.457241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T01:15:16.607346Z digest=sha256:a9c4bc3d29e27dfbff070f861f3005d1afa407087f3d4b2c91a9758c5eaf61ec

Observation c9232f9c-3633-4e5a-a66c-fc6f1252ba5f · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T06:07:41.216482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T12:55:47.754632Z digest=sha256:64dcf53d6d31654da741548f7eb19908ea022ad6cdb911d3d47404bfae6e3afa

Observation 160ffbf3-06dc-4a65-acd6-87ae5b677abf · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T18:07:09.018997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:07:09.018997Z digest=sha256:95be6b50dd03f5f57bea4c2c708a5a14d77214ed013543e5eb6fa2effd7be18b

Observation 5f2e8bd8-803a-481a-8d42-b2d78553f128 · inbound

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought cites this paper.

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:24:38.345154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T11:17:48.279808Z digest=sha256:894271e2e4a8697fe286349fd51b7a2e65bf634c35e93be8fc3321ef86d40c3a

Observation 935ccc88-2c86-48eb-9f65-6a23e5facef3 · inbound

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought cites this paper.

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T11:20:26.059257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:20:26.059257Z digest=sha256:6fa6c11b7a70cdb14daf33723c8cc4c3dffda0e83fde153df774383224fea3bf

Observation 937cde66-1be3-4a73-80ec-7cc7de59baca · inbound

Vesta: A Generalist Embodied Reasoning Model cites this paper.

Vesta: A Generalist Embodied Reasoning Model PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:29:35.801718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T16:55:12.518255Z digest=sha256:93c3276ff9e0bfd7a0594a4d91cb6125ee1167a859e69bc51c0b921a337b4169

Observation 62c9f909-28b2-4685-84fb-b53e16455f19 · inbound

Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction cites this paper.

Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:44:18.646376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T06:43:39.125052Z digest=sha256:d31951bc8fd3a1af472fbb307b37ecd0a01b30d2858cf87fd315d13061ad30fb

Observation eecd6d99-d4cc-4ace-8dbb-1fbaa235886a · inbound

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity cites this paper.

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:07:17.512742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-02T18:57:46.841456Z digest=sha256:8cfb0392100ab379b68db016bc64d654d71b117ef0d1cc2bbb35d6cbac967f57

Observation 6b4cab17-7f1f-4779-ac9d-03673e7b6a88 · inbound

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models cites this paper.

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T05:01:28.883339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T05:01:28.883339Z digest=sha256:19514e153d7380b0f3a888cfb332bb57cace4d8bf212402746957c2c66ee8d98