Pith. sign in

Paper Citation Record · LEDGER

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

As of 9 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 21 inbound Pith citation observations for arXiv:2505.15517.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15517 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:21:15.562904Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:56:37.191428Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:59:46.871127Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2810f855-b6ec-49ef-bff6-84e1f1ed0461 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Learning transferable visual models from natural language supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:10.661547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:10.661547Z digest=sha256:b04fbd70df1011a6df2cb945a5d4ce13e39e16316d1ceddb28d4cad6df8b2c23

Observation 46b306a3-2f10-4c8b-aeca-419db50ceaf2 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Qwen2.5: A party of foundation models, September 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:10.774245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:10.774245Z digest=sha256:1478e6fa9188ebe3dd7d37dd43b75a41a801ce02fdfbfe139ad2b9cd3b19252b

Observation a6f7afe0-db68-498b-9577-5941013cbed1 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:10.925490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:10.925490Z digest=sha256:a5c7cdf91a8dccc5b52e01f978b201bb10e9849250f15c864fa70627f932a219

Observation b283527a-94e2-4907-8815-859bfbf2d01c · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:10.987905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:10.987905Z digest=sha256:2f03747f0bebacac8c4e6e708d3cac4b74072a9b4c7880e22fe83a0bea3fa9fa

Observation c42254d0-0d1d-4b9c-8e6f-866eea83b1ce · outbound

This paper cites Claude 3.5 Sonnet.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Claude 3.5 Sonnet

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:21.793205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:11.067543Z digest=sha256:232c931d0a41a9d3f1dcec7ed677be142baf1395ed437f4b93178507e54d68de

Observation d9087d7c-72a7-4cd8-a609-f6258e7b940c · outbound

This paper cites GPT-4o System Card.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets GPT-4o System Card

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:21.645637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:11.157025Z digest=sha256:ddacee164cf74bfe5b18b88297833d29528e80df4b1795f1ce5a4274eee8bed8

Observation 88ad1e3c-4542-4184-9ad1-b9195e891e99 · outbound

This paper cites Gemini 2.5: Our most intelligent AI model.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Gemini 2.5: Our most intelligent AI model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:21.474603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:11.248985Z digest=sha256:a2ca2051e45f112b297c334581259f42e02c4653d08ea5af842341eff4d54095

Observation fc579134-c768-4c91-bc3b-cbe21cff9066 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:11.317851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:11.317851Z digest=sha256:380b788aa421793566e98848f1b72bb6eada3a1bc49e34946f452879a1844daa

Observation ef00d9c1-6dce-42c5-9560-73ebf9d79a41 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models, 2024.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Prismatic vlms: Investigating the design space of visually-conditioned language models, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:11.431731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:11.431731Z digest=sha256:b119e9bcae7db0df0944a0c964bfaa9c92570b80cffceae3d089982f9cac0fcf

Observation 5c840a04-665a-46c8-86d0-de2aaf072892 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets OpenVLA: An Open-Source Vision-Language-Action Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:11.524858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:11.524858Z digest=sha256:a939ecbb4e53e9fa8bf2926dc523f2496d5125bf303fe20cfe65605bf4169439

Observation 37ae5ddd-834d-4fde-8be4-6d2b444708e5 · outbound

This paper cites Gemini robotics: Bringing ai into the physical world, 2025.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Gemini robotics: Bringing ai into the physical world, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:21.320767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:11.617361Z digest=sha256:db1d5825f0b8ea6dcd6771182abd90caea59d273dc429f3eff84e66e9f7809b5

Observation e53a4628-f2d8-499f-a70b-27291e9247b3 · outbound

This paper cites Hi robot: Open-ended instruction following with hierarchical vision-language-action models, 2025.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Hi robot: Open-ended instruction following with hierarchical vision-language-action models, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:11.717767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:11.717767Z digest=sha256:0b87ea43e9340f3525761994be14c7131469edde9e11a73b58efefc1f5a3d07d

Observation 49756c7c-1796-49b5-ac41-6576e0eeaaac · outbound

This paper cites EQA-MX: Embodied question answering using multimodal expression.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets EQA-MX: Embodied question answering using multimodal expression

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:21.177361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:11.843051Z digest=sha256:15f3c21647466d0643aafab871563fd706e82494b4b7e6f468511a09739980a0

Observation d5486305-0b77-41c1-abe9-ad1f99cd4ad6 · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:11.971065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:11.971065Z digest=sha256:25b129bc7fd7daf5e095a770fd6dfa5b0fb242b6e1953b4c0d0e5101d066cf04

Observation a19ba6cb-a6b2-49fa-9a97-67f7ed454784 · outbound

This paper cites Embodied agent interface: Benchmarking LLMs for embodied decision making.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Embodied agent interface: Benchmarking LLMs for embodied decision making

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:20.976649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:12.071017Z digest=sha256:dd18f75e8427dbcf1c23b5af5c96da3f56882f551489c17ac9cc40259f70ddc4

Observation 0e4dabe6-a413-4a42-a05a-e30dad133706 · outbound

This paper cites ALFRED: A benchmark for interpreting grounded instructions for household robots.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets ALFRED: A benchmark for interpreting grounded instructions for household robots

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:20.803252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:12.204644Z digest=sha256:03cf9b12e34045184a02ba346da07751b30d69c70fabf40dbfca357e78af4070

Observation 4a5b6fda-a852-4c0e-87a3-444df8ed09ab · outbound

This paper cites Habitat 2.0: Training home assistants to rearrange their habitat.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Habitat 2.0: Training home assistants to rearrange their habitat

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:20.597082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:12.307005Z digest=sha256:82de6207a6cc2c1022b784119e43cfe61035f64f603c9b85c4bfef4552b5a2f1

Observation 89fac9a5-be2f-424e-96fa-2ecab99ce51e · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:12.433304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:12.433304Z digest=sha256:2a371e9404e910e95d7c4f361d7afd4a1f784c1cf36a44acb5e6800d79207108

Observation 28ffec7f-042f-4d23-a4be-d9ff5dd7e60b · outbound

This paper cites RoboVQA: Multimodal Long-Horizon Reasoning for Robotics.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets RoboVQA: Multimodal Long-Horizon Reasoning for Robotics

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:12.537196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:12.537196Z digest=sha256:b3c8f8e725605e0dc0e2391de9e477ebb8522e25a6c1518ec03a66c5b7057dab

Observation fbbefaac-f9b0-4556-84d0-13326ffaa9a0 · outbound

This paper cites Robobrain: A unified brain model for robotic manipulation from abstract to concrete.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Robobrain: A unified brain model for robotic manipulation from abstract to concrete

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:20.429334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:12.635029Z digest=sha256:adcfc7d8f5f9cddb4b916b34c4a859a70b76b792cf9f78d84d531969749cb272

Observation 8b530ecd-9b84-45c7-abe9-90bc9fa4d57f · outbound

This paper cites End-to-end training of deep visuomotor policies.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets End-to-end training of deep visuomotor policies

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:12.705074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:12.705074Z digest=sha256:a4333af5b810a55b9f1e7218246e22fb959256e28e5c6925c3955bf3dcc61f00

Observation b8bb77a2-732e-4716-a7e7-3ec37fe1b30c · outbound

This paper cites Octo: An open-source generalist robot policy.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Octo: An open-source generalist robot policy

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:12.758835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:12.758835Z digest=sha256:ff4b625e15b57d54f59ec94b60ca02723bac444d1a716d5ce198250537cc1644

Observation 7cb2f140-4df4-4fd6-8fd0-da4d723b5616 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Diffusion policy: Visuomotor policy learning via action diffusion

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:20.205356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:12.855366Z digest=sha256:58373f31be56480e73059e9d73e7700cb2f1836bdd0b6cc81557b6dfcb574ee7

Observation fadc7e27-88de-4086-9c27-191a49deb0b7 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:12.929203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:12.929203Z digest=sha256:5e3dcd7bf70bf5273568c2f5025c7bd8158a9e07c55c88d4373efd35706f9ec9

Observation 70eef0d6-9a22-4c5e-9eeb-66ec350dfdf6 · outbound

This paper cites Doremi: Optimizing data mixtures speeds up language model pretraining.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Doremi: Optimizing data mixtures speeds up language model pretraining

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:19.989185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:13.048192Z digest=sha256:2522e40c53c06e0008254bd90240d03e955ee4ff9b11545a8ea9cc713ffb85c0

Observation d8c58291-f7e2-482e-bbb5-1a69728bb5fa · outbound

This paper cites Remix: Optimizing data mixtures for large scale imitation learning.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Remix: Optimizing data mixtures for large scale imitation learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:19.661846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:13.138606Z digest=sha256:79a377db723c6707c32b29bcabd9705e8d3333e9b40d7be2b1f35d41deabca9f

Observation 570b1f98-6c35-4d47-b7b7-fefbbd8c99e2 · outbound

This paper cites Waveform Manipulation Against DNN-based Modulation Classification Attacks.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Waveform Manipulation Against DNN-based Modulation Classification Attacks

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T15:21:15.736057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:13.327764Z digest=sha256:204891266a635e4f67266ae4b0193886924569dac37b2f8a7173f3d65925cf63

Observation 677fab60-abfd-4446-93b7-ff132fdf478e · outbound

This paper cites Rt-1: Robotics transformer for real-world control at scale.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Rt-1: Robotics transformer for real-world control at scale

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:19.434336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:13.602732Z digest=sha256:14b0b09a1610a91f1cb0eafeeb198d93830f4ce65d105cbe4f86608db3a84bde

Observation 2536aae8-b720-4680-bf63-10254c34c0d8 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:19.132362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:13.799507Z digest=sha256:48e700141a1ea56aeaa4ed319c3c6ca0c9032ea6afab9856bab406dba88963e9

Observation f0380f50-e9a3-4b7f-9da8-cb92844a76ed · outbound

This paper cites https://physicalintelligence.company/blog/pi0, 2024.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets https://physicalintelligence.company/blog/pi0, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:18.775799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:13.879589Z digest=sha256:66f983fca27a35691e70b4af079117a01acdae77fcac48a2fa719c73fe3a4c2e

Observation 7139a414-069e-42f3-816c-a11d6f3d5eb5 · outbound

This paper cites Droid: A large-scale in-the-wild robot manipulation dataset.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Droid: A large-scale in-the-wild robot manipulation dataset

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:18.499528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:13.987722Z digest=sha256:8b99bb2df39ce98096d15e6aa57aab0c397061aa8e7e3dfb0ab61e87059703dd

Observation f0516040-c53b-4e48-83ed-da47be961464 · outbound

This paper cites Latent plans for task agnostic offline reinforcement learning.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Latent plans for task agnostic offline reinforcement learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:18.204100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:14.087953Z digest=sha256:a32ba40e378b3ddf27c1171c82331faf857a3d6a7b560a09ee5e1c37c175b043

Observation 4707bcc1-3f63-47b7-b916-875eb7225b2d · outbound

This paper cites Grounding language with visual affordances over unstructured data.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Grounding language with visual affordances over unstructured data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:14.185224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:14.185224Z digest=sha256:f2b01bcf86df33c454425a3ca8d9bde0f883f00bfb48ec3c5067b04dc71a3735

Observation 723865a2-2ef1-487b-850f-bfe2f1c56e39 · outbound

This paper cites Making sense of vision and touch: Self-supervised learning of multimodal representations for contact-rich tasks.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Making sense of vision and touch: Self-supervised learning of multimodal representations for contact-rich tasks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:17.966793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:14.254802Z digest=sha256:5b85fc17c5ec74b2590cc3817d7cc0fc55d3a8839a5cb6d115178022ab13ea88

Observation 721a2928-b8d1-4200-824a-7be393c81ebd · outbound

This paper cites Berkeley UR5 demonstration dataset.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Berkeley UR5 demonstration dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:14.336434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:14.336434Z digest=sha256:39c7eca594d182090b0229b69db0dbdf7fa50b940511aff5a3e6a33ae8efe02e

Observation 75e9d37e-1a9e-40d1-b3bd-f87294bb7b31 · outbound

This paper cites Robot learning on the job: Human-in-the-loop autonomy and learning during deployment.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Robot learning on the job: Human-in-the-loop autonomy and learning during deployment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:14.431315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:14.431315Z digest=sha256:8ab4ceb0f5f3a00ba884846203bd32b59ff4c5713161d6414d702d4ab8091f9c

Observation 3a37763b-cb4e-483e-bcb7-c240ec7642cd · outbound

This paper cites Real-world robot learning with masked visual pre-training.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Real-world robot learning with masked visual pre-training

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:17.727657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:14.499268Z digest=sha256:1b9a4935fba91f9121d2e79c0e2518091182909912ec00f5ef669251949af235

Observation 401421ae-7219-4c66-a579-481f77a0dbed · outbound

This paper cites The surprising effectiveness of representation learning for visual imitation, 2021.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets The surprising effectiveness of representation learning for visual imitation, 2021

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:17.481239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:14.577357Z digest=sha256:a83c6e9dc38b89e18768fef07dd3ab464cbc3279c27f34ac334d4f8c6bf3acf8

Observation 6e486d93-47b0-470d-92b1-8d69a030be3b · outbound

This paper cites Fanuc manipulation: A dataset for learning-based manipulation with fanuc mate 200id robot.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Fanuc manipulation: A dataset for learning-based manipulation with fanuc mate 200id robot

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:14.640236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:14.640236Z digest=sha256:7fdc8b885522b6519ef59ec3bdb70403e0453a9b7c7c224dca7dd710e1c81226

Observation f58f7461-6395-421a-b962-82b99b8a7214 · outbound

This paper cites Modularity through attention: Efficient training and transfer of language-conditioned policies for robot manipulation.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Modularity through attention: Efficient training and transfer of language-conditioned policies for robot manipulation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:17.381341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:14.737882Z digest=sha256:fe61509cdff91a82da8c83fe90f957910ba5d13f9a1f2aa7b9640e38df7463e3

Observation caa322b1-315a-47a2-8d3b-be8fcd050494 · outbound

This paper cites Learning modular language-conditioned robot policies through attention.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Learning modular language-conditioned robot policies through attention

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:17.152797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:14.840403Z digest=sha256:3fa3e653fc55a5669b26a35d1720c11a900c0cba07f57a51c8016cb377644611

Observation 3d21542a-2063-4be4-b112-479ab68a7519 · outbound

This paper cites Viola: Imitation learning for vision-based manipulation with object proposal priors.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Viola: Imitation learning for vision-based manipulation with object proposal priors

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:16.924497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:14.915362Z digest=sha256:0f1678fddc97fbf7da000d6bf0dc81f7f9e922bf823a93756c41f51b939d1116

Observation fada99ea-02d1-4d8e-afe3-83c6258a0d68 · outbound

This paper cites Bottom-up skill discovery from unsegmented demonstrations for long-horizon robot manipulation.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Bottom-up skill discovery from unsegmented demonstrations for long-horizon robot manipulation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:14.975665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:14.975665Z digest=sha256:8adc0f7c3be7602e8447e93b412d9cc1521aa0155413f19659c8f2592d76de0c

Observation b2c6f6d8-9800-4059-a84c-c4873317280e · outbound

This paper cites Watch and match: Supercharging imitation with regularized optimal transport.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Watch and match: Supercharging imitation with regularized optimal transport

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:16.689000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:15.034438Z digest=sha256:8d5088e02b7f9f36b095693218269443f8dfadcb1c1516e1a7b96f69855e6954

Observation 2c736d5f-51b0-4d83-b430-8bb752612624 · outbound

This paper cites Vqa: Visual question answering.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Vqa: Visual question answering

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:15.093700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:15.093700Z digest=sha256:d052b0f94ded59c9838e3b91d53efc8d0fec8f19fa1f07f93117d4baed3f9bbd

Observation 8d6a418f-6093-407c-94c2-7a97097d269f · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:15.171551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:15.171551Z digest=sha256:bdb3ef9eec703087c4f0d290af5f6784c70d00cf25a74de6b5c6c17f52beeac3

Observation 9a429751-fb69-4ff3-b70e-16c03ab4fd54 · outbound

This paper cites What is Right for Me is Not Yet Right for You: A Dataset for Grounding Relative Directions via Multi-Task Learning.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets What is Right for Me is Not Yet Right for You: A Dataset for Grounding Relative Directions via Multi-Task Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:15.234679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:15.234679Z digest=sha256:7fcf14052670417c285132e820242050705a4bb527ab47b58d0c2acf6abc3609

Observation aa95f25d-8413-4f1c-bdc6-dff3c0d17766 · outbound

This paper cites Embodied question answering.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Embodied question answering

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:16.492098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:15.307972Z digest=sha256:e989a4b5444acf87515c5628a1170aa24c436afdcd5795c9ccece580a242a044

Observation d9b0369e-a073-456f-96b9-c14cb3d243d3 · outbound

This paper cites Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:16.324862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:15.362252Z digest=sha256:a10c078794d4edf8a0d1995a35ec88376abd0fedc6ecba89d9914be158ba4821

Observation 2320ef47-b91b-4d3a-a440-3a5a62ef015f · outbound

This paper cites Christensen and Gregory D.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Christensen and Gregory D

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:16.140182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:15.406161Z digest=sha256:0f4335ef392cbd8e984e120748dcd62a7c679d83ae784a286becb7f1cca0a4e6

Observation 8c0dbc78-7d6e-4cc4-8b85-fdebb3c0ba27 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:15.503226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:15.503226Z digest=sha256:babb11553e35b0a5bac849d67ca78345bd808d6c44f4ff192d26b6f05e37fa30

Observation 323bc312-f722-403f-a991-2854cc6ed622 · outbound

This paper cites Configuration D.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Configuration D

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:15.955382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:21:15.562904Z digest=sha256:a086113befd00916c3cb7585b6fd36ae125e38631bdf6b4d91efbd00d9261c0a

Pith citing papers

Observation 3958f27a-cd4c-4d43-a2c9-5896bb4b176d · inbound

Demystifying the Visual Quality Paradox in Multimodal Large Language Models cites this paper.

Demystifying the Visual Quality Paradox in Multimodal Large Language Models Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:56:37.191428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:56:37.191428Z digest=sha256:8718a3ecb1b816af652d34e80d804a54ad1a33bfa3909e785268667b71428720

Observation 71e0cde9-4a6e-48b7-bf08-c46222e1b4e5 · inbound

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction cites this paper.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:07.721472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:07.721472Z digest=sha256:f1bd6e979a2633caea6ab75f617a4dcb682b4a50550a74991166a55fd8063a31

Observation b5371277-8e79-4797-ab8d-c1eabf6250ec · inbound

Foundation Model Driven Robotics: A Comprehensive Review cites this paper.

Foundation Model Driven Robotics: A Comprehensive Review Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T17:43:53.200688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:43:53.200688Z digest=sha256:2d17c5fe628d377763a6cbb9962656d8d3b049be8fc3d403a7c210a4fa6be432

Observation 84347f74-b78d-436e-b668-2d701380c58b · inbound

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training cites this paper.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T13:29:54.177769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:29:54.177769Z digest=sha256:e09d36e98b4e728020158090f8604e2699811a270fb247dec61a71b68ce64535

Observation ab07615c-fe48-4386-9ebc-e6962892d043 · inbound

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning cites this paper.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:6ff10fec19b2c36825cdce4f105c49c60a31920ec60a5a1ce8b72a28cf84842d

Observation db4d926e-04c7-4446-82ee-35e3b02bbd32 · inbound

EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training cites this paper.

EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:11:03.769973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T02:16:08.687340Z digest=sha256:4c11bb526b665ebd1c1bbda0838be49d9e8e6a65895d746797aa63ec2bfa4918

Observation 7d10b362-674a-4162-8a8b-2ffdfcf7c6aa · inbound

Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models cites this paper.

Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:37:35.557834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T18:35:22.595183Z digest=sha256:664e727fdbd7ac550928d25f852efa14b5beade47b23ca8224076a6c5531dad8

Observation 6597ac9a-428c-4f63-81ed-4b8b7612e738 · inbound

RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents cites this paper.

RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:08:05.177944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:05:42.451746Z digest=sha256:fba6d8056f02b78b9ef75441b769dba1a4ee4035637c19a19f73f1d41114b3e9

Observation 13d6f012-16b6-44f1-a40d-981388de4813 · inbound

Rethinking VLM Representation for VLA Initialization cites this paper.

Rethinking VLM Representation for VLA Initialization Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:24:00.128773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T22:21:29.733181Z digest=sha256:3497ccd7faf3dccf08d64cbc3e7d81a354b3c60bff9367bbd660cc64e616af84

Observation 69813e01-6f70-4049-8b54-9152fed034f7 · inbound

Extending Embodied Question Answering from Perception to Decision cites this paper.

Extending Embodied Question Answering from Perception to Decision Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T21:33:58.981062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T21:30:40.182958Z digest=sha256:6f2f0bc95be10b1ca6c6611d1cf34bf14a59b6e14b648f95f0df4d886df8ec3f

Observation 78ecdf25-e94e-4225-b0d2-a12583a15c8c · inbound

GEM: Generative Supervision Helps Embodied Intelligence cites this paper.

GEM: Generative Supervision Helps Embodied Intelligence Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:43:28.939856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:3e86254780993b31ebbc62773c462ada2c422bbebe50f54cf0a9562f6e5c4154

Observation 6657bccb-16d3-4b2f-b549-c6a75779ff99 · inbound

Wall-OSS-0.5 Technical Report cites this paper.

Wall-OSS-0.5 Technical Report Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:26:00.530454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T22:35:00.258436Z digest=sha256:966589ae61e3bb3bc216a44848697f8fc2f729435c8d4a18523cf585973f6d4a

Observation 6d129ccf-4039-44f3-a1d0-427ab9ba868a · inbound

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data cites this paper.

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.630648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T18:31:03.548002Z digest=sha256:c8511192571b1594dcdc9356f2555f23724d6f7b230e19368581079c804e396f

Observation ea05806d-e451-4a73-87e4-ebb4d88fd28c · inbound

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation cites this paper.

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:34.770570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T06:26:32.209719Z digest=sha256:29d655ba447a987ace3ddf5c3c3a973deb31ef9e8675a495a65d3962e483104f

Observation 93a450d5-0507-48c7-8596-c619c4fb2a52 · inbound

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation cites this paper.

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T11:43:49.379947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:43:49.379947Z digest=sha256:2c5b0a1bab0e62788cff8e249998566731cb605273610e502993db2446875329

Observation e9149a18-0ed3-4570-b74c-381e2a561269 · inbound

Vesta: A Generalist Embodied Reasoning Model cites this paper.

Vesta: A Generalist Embodied Reasoning Model Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:29:35.736977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T16:55:12.518255Z digest=sha256:e0eac5eb619a38506b1ec9358c12270ed0503e3f0a2ff2ecfa7fec2b447b7f35

Observation 974cc6a3-9b1b-4118-884a-d95e485da3ce · inbound

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification cites this paper.

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:59:46.873792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T08:12:14.829556Z digest=sha256:295fe00c43a7c7e9d0bffdafa3a252a5ccc8e2017cdae0fa0ac08abfd51f279d

Observation a99f9c32-9729-45bf-81f1-806a6e7e4970 · inbound

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos cites this paper.

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-14T17:30:54.988498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:30:54.988498Z digest=sha256:b1a449579aa750515487c3ee685e7f5879e838bdcf1413937bee037750f595fb

Observation f7c34079-fe85-4677-91b1-0d98d1ed7f7e · inbound

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment cites this paper.

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T05:16:35.354226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:16:35.354226Z digest=sha256:53547b09aa21ca85e6f1006018fb9b7b22ac0ca29aae6c8430f399816adc7a88

Observation 0ac0bafc-1318-471d-8f61-4b03a1718018 · inbound

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model cites this paper.

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T16:32:46.402572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:32:46.402572Z digest=sha256:d8676b75a570581562db3f6d8bdf53d934bd0384570377407c50dc93bcb9cd39

Observation 21826f6d-47ea-4d4c-be0f-410311188457 · inbound

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model cites this paper.

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T01:57:27.437916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:57:27.437916Z digest=sha256:8839d4743dfd3fa4c4935643d88893ce2c4691e547999808f361ae781c8424b7