Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:24:01.169260Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 16 inbound Pith citation observations for arXiv:2505.09990.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:24:01.169260Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T11:20:26.059257Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T04:29:35.800189Z
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2219b9ac-df78-42ae-9265-3b8a45fb70c8 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dee1d29d-305f-471b-b72e-f55f77c9c521 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9ad60340-6412-4c9f-b267-5d6fb77174b0 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing The claude 3 model family: Opus, sonnet, haiku
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6dc0f7fa-e12b-442d-ad87-56222dca88da · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Qwen2.5-VL Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cda1af7f-3927-41c0-88ad-921203d727a5 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b51eda-55be-4b02-a2fa-93d594d6343e · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Vizwiz: nearly real-time answers to visual questions
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8449394f-2901-494b-b8e9-1541c3e67514 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 632fb135-1918-402e-a31a-5a11440682d5 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9be612a2-71b3-4ef0-b76f-7e313b9d82ff · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing GuessWhat?! Visual object discovery through multi-modal dialogue
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf4f6c35-192a-4d9f-8515-eea40efbd1dd · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3862487-399c-48ed-a95d-aeece23da61e · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing AR2-D2:Training a Robot Without a Robot
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63440c4a-c670-41f1-904c-b6789f1fb571 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f813e39e-b487-4ea8-95d2-fbbf9ec92f9f · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d706436b-bbe4-47d2-a168-f184c2000066 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c913113c-306a-4268-b462-da67fe257c52 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d9ca603-53b9-4985-b0f9-785d1f42c736 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Referitgame: Referring to objects in photographs of natural scenes
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ebfc7ee-b121-4fa5-9a4a-7f1e487e0763 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Tournament evaluation of large language models, 2025
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 11a55da1-0788-4a55-8ef5-9665e2b508a7 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Segment Anything
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54b65a26-b2a7-4707-8519-e3abb3f8ca96 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing LLaVA-OneVision: Easy Visual Task Transfer
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12678563-b8c7-4c75-8dd0-f305f3bdb733 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing am-ELO: A Stable Framework for Arena-based LLM Evaluation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b74a53f2-a5f5-4ddc-b1b6-c0683c186ad7 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Improving Your Model Ranking on Chatbot Arena by Vote Rigging
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7869ea04-7d41-4517-8449-17f882bea05e · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing CityRefer: Geography-aware 3D Visual Grounding Dataset on City-scale Point Cloud Data
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec0e5f45-3049-49ce-b7bd-a866edb56ba4 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0a08dfc-a950-42ab-99cc-707b47749c09 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Multimodal Explanations: Justifying Decisions and Pointing to the Evidence
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 878890d2-951f-4555-9dd9-6a3e351a436f · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a35e52b8-ba94-426b-ad5e-8fbaf3dc8032 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Sat: Spatial aptitude training for multimodal language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4cb9e7e-9e74-40e8-8ec7-66cd1e6ccd70 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Openarena: An open platform for llm-as-a-judge evaluation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 66b74e56-b0f8-401c-a958-bd3afd85dd1c · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Gemini Robotics: Bringing AI into the Physical World
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38c0aae0-0a95-45c3-9736-81b688d373e1 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing A new look at infant pointing
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b1e0e513-6bbe-46e9-8132-81fb3f4c6289 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing A Comprehensive Review of Multimodal Large Language Models: Performance and Challenges Across Different Tasks
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afaaef75-58a1-4208-b2f1-c4ca25ba31a8 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Chain-of-thought prompting elicits reasoning in large language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd9ccb08-3e39-43a0-9298-1f57216104a4 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Grok-2 model card
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0a5b8dcc-b802-46d4-a45d-f2ab9418877f · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing A survey on multimodal large language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f848a5d-4384-4d58-80ae-137b1b0428e9 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Modeling Context in Referring Expressions
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1630d39-e1d2-42ae-9e1e-1d2b95845a11 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d42ccffb-65a6-4b6b-b42a-0822ff0a10db · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Robopoint: A vision-language model for spatial affordance prediction in robotics
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c6da27ed-7ce2-45c4-888c-e085982500bb · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe170848-47af-48fe-8b08-2a2d80520cb2 · outbound
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73db2bcf-1a51-4c74-a4f5-8d545ab2785a · inbound
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation baddc7ca-87d9-45e0-86da-474f4af62f9c · inbound
Seed1.8 Model Card: Towards Generalized Real-World Agency PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3ce27e9b-59bd-4986-a77a-a854a49b07d5 · inbound
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b7872288-c099-44fe-af46-6b1937304fba · inbound
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ccd375ba-116a-4bfe-b769-a790e05b02b3 · inbound
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8d918e7f-bd15-4dbe-bdea-45df91653455 · inbound
MolmoAct2: Action Reasoning Models for Real-world Deployment PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9fc55664-f5a3-49a2-9a06-8e511421e035 · inbound
MolmoAct2: Action Reasoning Models for Real-world Deployment PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 85f87663-0c8f-42a6-b2dd-83a67e1e51be · inbound
ZAYA1-VL-8B Technical Report PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Reference 146
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c9232f9c-3633-4e5a-a66c-fc6f1252ba5f · inbound
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 160ffbf3-06dc-4a65-acd6-87ae5b677abf · inbound
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f2e8bd8-803a-481a-8d42-b2d78553f128 · inbound
RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 935ccc88-2c86-48eb-9f65-6a23e5facef3 · inbound
RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 937cde66-1be3-4a73-80ec-7cc7de59baca · inbound
Vesta: A Generalist Embodied Reasoning Model PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 62c9f909-28b2-4685-84fb-b53e16455f19 · inbound
Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation eecd6d99-d4cc-4ace-8dbb-1fbaa235886a · inbound
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6b4cab17-7f1f-4779-ac9d-03673e7b6a88 · inbound
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.