Pith. sign in

Paper Citation Record · LEDGER

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

As of 17 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 22 inbound Pith citation observations for arXiv:2505.15517.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15517 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:21:15.562904Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:18:34.179159Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:59:46.871127Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2810f855-b6ec-49ef-bff6-84e1f1ed0461 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Learning transferable visual models from natural language supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:10.661547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:10.661547Z digest=sha256:a1c2c8b3546a137a46ed5d5ca6f6c55278e006cf946293040d91e2810f9ec513

Observation 46b306a3-2f10-4c8b-aeca-419db50ceaf2 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Qwen2.5: A party of foundation models, September 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:10.774245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:10.774245Z digest=sha256:801a0d9ef3005bfb93bd6faf3c0007efdb6639265c2983a9f213ec434a83d13d

Observation a6f7afe0-db68-498b-9577-5941013cbed1 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:10.925490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:10.925490Z digest=sha256:3e1142f6004ed75a6818bd98f1d73ab15555dea49ededd52c1c711ede1971fb9

Observation b283527a-94e2-4907-8815-859bfbf2d01c · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:10.987905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:10.987905Z digest=sha256:8ac19dc52e088beaf0c318ae023b46bd428b56d1472ba9269a9227306e6f429d

Observation c42254d0-0d1d-4b9c-8e6f-866eea83b1ce · outbound

This paper cites Claude 3.5 Sonnet.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Claude 3.5 Sonnet

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:21.793205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:11.067543Z digest=sha256:148651a1b09df183047126c31c3d3d01d09ccf67bf5eef4155c856df29ce4ae3

Observation d9087d7c-72a7-4cd8-a609-f6258e7b940c · outbound

This paper cites GPT-4o System Card.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets GPT-4o System Card

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:21.645637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:11.157025Z digest=sha256:8f59c7cf7d56efac4a4fe7818a942fe94517efc1e08b2f73f8fd25af52222dc4

Observation 88ad1e3c-4542-4184-9ad1-b9195e891e99 · outbound

This paper cites Gemini 2.5: Our most intelligent AI model.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Gemini 2.5: Our most intelligent AI model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:21.474603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:11.248985Z digest=sha256:442edf9b00b1b228bf855b7baa43d2201637cc85834d143b3d85b4d93e04c968

Observation fc579134-c768-4c91-bc3b-cbe21cff9066 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:11.317851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:11.317851Z digest=sha256:4cfe3f7c7857c5c12390bfab98d5ccc614f26d65ec6f19739705c3e2f0938854

Observation ef00d9c1-6dce-42c5-9560-73ebf9d79a41 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models, 2024.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Prismatic vlms: Investigating the design space of visually-conditioned language models, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:11.431731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:11.431731Z digest=sha256:53852746a23b5aa0a3f2c0676d46ebe33e4453aa66ec85e60248cfffd5ead6c4

Observation 5c840a04-665a-46c8-86d0-de2aaf072892 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets OpenVLA: An Open-Source Vision-Language-Action Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:11.524858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:11.524858Z digest=sha256:c8c2c498c30eee41b879847904114e7b8d4866a8330f29ade1bfe452bcdd71bb

Observation 37ae5ddd-834d-4fde-8be4-6d2b444708e5 · outbound

This paper cites Gemini robotics: Bringing ai into the physical world, 2025.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Gemini robotics: Bringing ai into the physical world, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:21.320767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:11.617361Z digest=sha256:39b7819d159aac0675b9ca5d5585114ec0effe4b67e43258c65986ffa6f7f8a7

Observation e53a4628-f2d8-499f-a70b-27291e9247b3 · outbound

This paper cites Hi robot: Open-ended instruction following with hierarchical vision-language-action models, 2025.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Hi robot: Open-ended instruction following with hierarchical vision-language-action models, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:11.717767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:11.717767Z digest=sha256:f7498277fdacd5a46b4aa2a15037eb30fbc50d658fde08cbfaae8397741d84d4

Observation 49756c7c-1796-49b5-ac41-6576e0eeaaac · outbound

This paper cites EQA-MX: Embodied question answering using multimodal expression.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets EQA-MX: Embodied question answering using multimodal expression

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:21.177361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:11.843051Z digest=sha256:b8075874827c92d87ba0b01a62a8c998b79f2baed63f84dba14ad61784b6a46e

Observation d5486305-0b77-41c1-abe9-ad1f99cd4ad6 · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:11.971065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:11.971065Z digest=sha256:8c4514d82387b58be59f0129a4415b39957a291ffd1028488eb6c45024889d07

Observation a19ba6cb-a6b2-49fa-9a97-67f7ed454784 · outbound

This paper cites Embodied agent interface: Benchmarking LLMs for embodied decision making.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Embodied agent interface: Benchmarking LLMs for embodied decision making

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:20.976649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:12.071017Z digest=sha256:255dd08ac57607ef6586a66fc900457d08b1df6278a473302b09e9d645fa7c20

Observation 0e4dabe6-a413-4a42-a05a-e30dad133706 · outbound

This paper cites ALFRED: A benchmark for interpreting grounded instructions for household robots.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets ALFRED: A benchmark for interpreting grounded instructions for household robots

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:20.803252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:12.204644Z digest=sha256:ad4aa20f838843f6b9f28670dd33b28f22714cb54aa5f90eec8dce0e832572d4

Observation 4a5b6fda-a852-4c0e-87a3-444df8ed09ab · outbound

This paper cites Habitat 2.0: Training home assistants to rearrange their habitat.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Habitat 2.0: Training home assistants to rearrange their habitat

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:20.597082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:12.307005Z digest=sha256:f79d57454a50ceb3515515f556a737c10c4d0bead805c2f4717429660a79e15f

Observation 89fac9a5-be2f-424e-96fa-2ecab99ce51e · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:12.433304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:12.433304Z digest=sha256:d730cd0104855955ee5b1567ea4d4da3a2e0be83056e7f7f731219a13918c07e

Observation 28ffec7f-042f-4d23-a4be-d9ff5dd7e60b · outbound

This paper cites RoboVQA: Multimodal Long-Horizon Reasoning for Robotics.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets RoboVQA: Multimodal Long-Horizon Reasoning for Robotics

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:12.537196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:12.537196Z digest=sha256:cdd942906510b8b9c7ee2d53105b8918760bf69a739abdca1d8d902666c95e52

Observation fbbefaac-f9b0-4556-84d0-13326ffaa9a0 · outbound

This paper cites Robobrain: A unified brain model for robotic manipulation from abstract to concrete.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Robobrain: A unified brain model for robotic manipulation from abstract to concrete

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:20.429334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:12.635029Z digest=sha256:8181ff9f68fcb36946c94feca86042258cf442e309d24637e03693570cfe48d4

Observation 8b530ecd-9b84-45c7-abe9-90bc9fa4d57f · outbound

This paper cites End-to-end training of deep visuomotor policies.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets End-to-end training of deep visuomotor policies

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:12.705074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:12.705074Z digest=sha256:a039f331ce60c76aba97478bdfee797969a43ced06eacac0073030eda1cbf520

Observation b8bb77a2-732e-4716-a7e7-3ec37fe1b30c · outbound

This paper cites Octo: An open-source generalist robot policy.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Octo: An open-source generalist robot policy

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:12.758835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:12.758835Z digest=sha256:85f64dd0313a320c96bdd7cae61932bae73ed67774e93ab6bef3a674278cd968

Observation 7cb2f140-4df4-4fd6-8fd0-da4d723b5616 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Diffusion policy: Visuomotor policy learning via action diffusion

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:20.205356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:12.855366Z digest=sha256:c876f7219b530a2a9d2a037245db1fe9796463045f442a5e6fb38d161df28d5d

Observation fadc7e27-88de-4086-9c27-191a49deb0b7 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:12.929203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:12.929203Z digest=sha256:7d495c5c3f01e1dcf507c82136ad13a4174e2b29f6b4076cc6f37179c30253ba

Observation 70eef0d6-9a22-4c5e-9eeb-66ec350dfdf6 · outbound

This paper cites Doremi: Optimizing data mixtures speeds up language model pretraining.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Doremi: Optimizing data mixtures speeds up language model pretraining

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:19.989185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:13.048192Z digest=sha256:4350e6da83e887a8390818bd2a910bdd6980719ca03d64805a4ec2ff24364b80

Observation d8c58291-f7e2-482e-bbb5-1a69728bb5fa · outbound

This paper cites Remix: Optimizing data mixtures for large scale imitation learning.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Remix: Optimizing data mixtures for large scale imitation learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:19.661846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:13.138606Z digest=sha256:156ca30e2e1d6404df170eb055df967e6867b75155a923eaa0857d56c6665b21

Observation 570b1f98-6c35-4d47-b7b7-fefbbd8c99e2 · outbound

This paper cites Waveform Manipulation Against DNN-based Modulation Classification Attacks.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Waveform Manipulation Against DNN-based Modulation Classification Attacks

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T15:21:15.736057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:13.327764Z digest=sha256:6dad51ecd1fb84943560e5c5d99856c0ee707f4929c65b64a5429307b048810d

Observation 677fab60-abfd-4446-93b7-ff132fdf478e · outbound

This paper cites Rt-1: Robotics transformer for real-world control at scale.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Rt-1: Robotics transformer for real-world control at scale

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:19.434336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:13.602732Z digest=sha256:9672dcd552efb298aad80065d3794a446608ad2bbff3ad603f777311464c8ee9

Observation 2536aae8-b720-4680-bf63-10254c34c0d8 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:19.132362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:13.799507Z digest=sha256:a949a73ec6f6a266b6d76b73b60c16d0f76e773d5e7c9e6d1503bafe05c373c9

Observation f0380f50-e9a3-4b7f-9da8-cb92844a76ed · outbound

This paper cites https://physicalintelligence.company/blog/pi0, 2024.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets https://physicalintelligence.company/blog/pi0, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:18.775799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:13.879589Z digest=sha256:5beab86c9fd7899a2caa9c4b723487cff04a530f2c8889f50f86d89a3ac463e4

Observation 7139a414-069e-42f3-816c-a11d6f3d5eb5 · outbound

This paper cites Droid: A large-scale in-the-wild robot manipulation dataset.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Droid: A large-scale in-the-wild robot manipulation dataset

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:18.499528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:13.987722Z digest=sha256:14793a337d1de76130c9402d108000a8f3bd22d04119f883af5f8c910402f99a

Observation f0516040-c53b-4e48-83ed-da47be961464 · outbound

This paper cites Latent plans for task agnostic offline reinforcement learning.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Latent plans for task agnostic offline reinforcement learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:18.204100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:14.087953Z digest=sha256:41a71c637a7de062adfb2b7124d50a84cc12ff44cc2ff026f6d0532f1ff54c60

Observation 4707bcc1-3f63-47b7-b916-875eb7225b2d · outbound

This paper cites Grounding language with visual affordances over unstructured data.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Grounding language with visual affordances over unstructured data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:14.185224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:14.185224Z digest=sha256:af463982355442ef55e617b531506a8fada1e6e49898b846406fb838fc2f455c

Observation 723865a2-2ef1-487b-850f-bfe2f1c56e39 · outbound

This paper cites Making sense of vision and touch: Self-supervised learning of multimodal representations for contact-rich tasks.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Making sense of vision and touch: Self-supervised learning of multimodal representations for contact-rich tasks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:17.966793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:14.254802Z digest=sha256:fd6e5e2515199b1d48ea5c4b1601131ac18fc578373e8e371c188cd3f3079f46

Observation 721a2928-b8d1-4200-824a-7be393c81ebd · outbound

This paper cites Berkeley UR5 demonstration dataset.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Berkeley UR5 demonstration dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:14.336434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:14.336434Z digest=sha256:952b04196aa9eb9bc5b67e273cae80e5dceb1461ce6de868962aac7af8629ed4

Observation 75e9d37e-1a9e-40d1-b3bd-f87294bb7b31 · outbound

This paper cites Robot learning on the job: Human-in-the-loop autonomy and learning during deployment.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Robot learning on the job: Human-in-the-loop autonomy and learning during deployment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:14.431315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:14.431315Z digest=sha256:a8ab90f48d9eab9468ba27c5c23cc25555b7701c263d5abc4bb44a93b26e8624

Observation 3a37763b-cb4e-483e-bcb7-c240ec7642cd · outbound

This paper cites Real-world robot learning with masked visual pre-training.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Real-world robot learning with masked visual pre-training

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:17.727657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:14.499268Z digest=sha256:dcf0649114c2b0e62ae9e1470896a067de6b96530ce2bb2c06e264598af5ed5e

Observation 401421ae-7219-4c66-a579-481f77a0dbed · outbound

This paper cites The surprising effectiveness of representation learning for visual imitation, 2021.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets The surprising effectiveness of representation learning for visual imitation, 2021

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:17.481239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:14.577357Z digest=sha256:761be8dd6468fc86dac17674b0ec8d1d7b2e1de4ba1959aac82a1d12ab3d70ab

Observation 6e486d93-47b0-470d-92b1-8d69a030be3b · outbound

This paper cites Fanuc manipulation: A dataset for learning-based manipulation with fanuc mate 200id robot.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Fanuc manipulation: A dataset for learning-based manipulation with fanuc mate 200id robot

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:14.640236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:14.640236Z digest=sha256:bf5b94c4e108f8c2f39ff5622447d0bfc097d41eda5af393c5fbf90efb80e876

Observation f58f7461-6395-421a-b962-82b99b8a7214 · outbound

This paper cites Modularity through attention: Efficient training and transfer of language-conditioned policies for robot manipulation.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Modularity through attention: Efficient training and transfer of language-conditioned policies for robot manipulation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:17.381341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:14.737882Z digest=sha256:ff44bbb7f9341ab061b9dfc3ec4a76a5c07786f497b0ec2faf89f3a2fae45354

Observation caa322b1-315a-47a2-8d3b-be8fcd050494 · outbound

This paper cites Learning modular language-conditioned robot policies through attention.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Learning modular language-conditioned robot policies through attention

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:17.152797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:14.840403Z digest=sha256:96b20ff1b370881119102c48c07e98ca9da010b130ffe0a67a8cbdbfaac6fff8

Observation 3d21542a-2063-4be4-b112-479ab68a7519 · outbound

This paper cites Viola: Imitation learning for vision-based manipulation with object proposal priors.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Viola: Imitation learning for vision-based manipulation with object proposal priors

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:16.924497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:14.915362Z digest=sha256:b0db8787855389ef84491766b119e921826b7f833086cb0a7a40521e3a5b80db

Observation fada99ea-02d1-4d8e-afe3-83c6258a0d68 · outbound

This paper cites Bottom-up skill discovery from unsegmented demonstrations for long-horizon robot manipulation.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Bottom-up skill discovery from unsegmented demonstrations for long-horizon robot manipulation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:14.975665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:14.975665Z digest=sha256:65ea244a94a42a2294797b56b1766d5297035d753f230357a210cd4cf5149a6e

Observation b2c6f6d8-9800-4059-a84c-c4873317280e · outbound

This paper cites Watch and match: Supercharging imitation with regularized optimal transport.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Watch and match: Supercharging imitation with regularized optimal transport

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:16.689000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:15.034438Z digest=sha256:734ab59db3d5014f9150d5e0229485705bf73f666f2c08c07391709459585c6f

Observation 2c736d5f-51b0-4d83-b430-8bb752612624 · outbound

This paper cites Vqa: Visual question answering.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Vqa: Visual question answering

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:15.093700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:15.093700Z digest=sha256:3276f2fdd0766fee22bf975cca18b69d123f4bf26d22ed5651cf5ccbad297c58

Observation 8d6a418f-6093-407c-94c2-7a97097d269f · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:15.171551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:15.171551Z digest=sha256:97ee850eaea1a85e32ca700ac665577cd8cdfb2f08a8046ef1640adef84db529

Observation 9a429751-fb69-4ff3-b70e-16c03ab4fd54 · outbound

This paper cites What is Right for Me is Not Yet Right for You: A Dataset for Grounding Relative Directions via Multi-Task Learning.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets What is Right for Me is Not Yet Right for You: A Dataset for Grounding Relative Directions via Multi-Task Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:15.234679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:15.234679Z digest=sha256:9c200627d1f507ebd1b3964befd6458bacf53bed575d650b113ac2aed10d9dd9

Observation aa95f25d-8413-4f1c-bdc6-dff3c0d17766 · outbound

This paper cites Embodied question answering.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Embodied question answering

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:16.492098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:15.307972Z digest=sha256:50d3b874dfe8a976d9bb0fc035023e3b30dd33d975663b3dddb1aaadc8c0f3d8

Observation d9b0369e-a073-456f-96b9-c14cb3d243d3 · outbound

This paper cites Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:16.324862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:15.362252Z digest=sha256:fb594a2a62352e3562e756affe365700e257615135def9b14565a17d967aee2e

Observation 2320ef47-b91b-4d3a-a440-3a5a62ef015f · outbound

This paper cites Christensen and Gregory D.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Christensen and Gregory D

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:16.140182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:15.406161Z digest=sha256:3bd28b63a992358b803a47b2d6dbd0a2f22f3b282969aab3db9368e8e2c4a7a2

Observation 8c0dbc78-7d6e-4cc4-8b85-fdebb3c0ba27 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:15.503226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:21:15.503226Z digest=sha256:a83404a05a17579a3cc87bc7c427210f3f16860a41bb2eebc9bd936bf1a886bb

Observation 323bc312-f722-403f-a991-2854cc6ed622 · outbound

This paper cites Configuration D.

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Configuration D

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:15.955382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:21:15.562904Z digest=sha256:4519b5dd32ac34a13dbb8198967b00a3eeb0505f47a7eebde6120ccdfa3adacb

Pith citing papers

Observation 3958f27a-cd4c-4d43-a2c9-5896bb4b176d · inbound

Demystifying the Visual Quality Paradox in Multimodal Large Language Models cites this paper.

Demystifying the Visual Quality Paradox in Multimodal Large Language Models Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:56:37.191428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:56:37.191428Z digest=sha256:a389fbe47fe11a7ea47634f5c6af434d6f57abb37b335d359d6744ebab71914e

Observation 71e0cde9-4a6e-48b7-bf08-c46222e1b4e5 · inbound

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction cites this paper.

HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:07.721472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:07.721472Z digest=sha256:b294976f5ea7da6a607eaa065e3aa7cd9f69a5b3234b8d2f69e19cff06b6acf6

Observation b5371277-8e79-4797-ab8d-c1eabf6250ec · inbound

Foundation Model Driven Robotics: A Comprehensive Review cites this paper.

Foundation Model Driven Robotics: A Comprehensive Review Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T17:43:53.200688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:43:53.200688Z digest=sha256:509c82848163ec7da40d9aa05c07bc8f51a82e6626c9a917ffb4c9ed443a25a1

Observation 84347f74-b78d-436e-b668-2d701380c58b · inbound

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training cites this paper.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T13:29:54.177769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:29:54.177769Z digest=sha256:75feaa79dd1e6149716e337a73550ea471819ae74ca6d25cf6188393a6c6673a

Observation ab07615c-fe48-4386-9ebc-e6962892d043 · inbound

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning cites this paper.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:413315549cf3523892bbbf46de7faab5ec5f6fbc5266420d2375e33164e78e0d

Observation db4d926e-04c7-4446-82ee-35e3b02bbd32 · inbound

EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training cites this paper.

EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:11:03.769973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T02:16:08.687340Z digest=sha256:d534782fbef2a42180991ad1b3c93cd07fd71ac0ce88a5e8c8a9be60e0466149

Observation 7d10b362-674a-4162-8a8b-2ffdfcf7c6aa · inbound

Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models cites this paper.

Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:37:35.557834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T18:35:22.595183Z digest=sha256:33977d29815a9b55e5c0a5691068f19eaa26fa77524c3f2742cb4c7a70f05a16

Observation 6597ac9a-428c-4f63-81ed-4b8b7612e738 · inbound

RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents cites this paper.

RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:08:05.177944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:05:42.451746Z digest=sha256:ba2956b8d79a0a7bcbbfd0bfbfc550f98c1b8cb10597fe714fca0296d8f33cd1

Observation 13d6f012-16b6-44f1-a40d-981388de4813 · inbound

Rethinking VLM Representation for VLA Initialization cites this paper.

Rethinking VLM Representation for VLA Initialization Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:24:00.128773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T22:21:29.733181Z digest=sha256:35d3901d43fd0dfc078de55f2107cf2e6b2b77cdfe00a0784d3fc7e2161484ba

Observation 69813e01-6f70-4049-8b54-9152fed034f7 · inbound

Extending Embodied Question Answering from Perception to Decision cites this paper.

Extending Embodied Question Answering from Perception to Decision Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T21:33:58.981062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T21:30:40.182958Z digest=sha256:4de1f233530b4ada8393c0c78b254037abfe5e58176dc4b7014ba76d7b7bceac

Observation 78ecdf25-e94e-4225-b0d2-a12583a15c8c · inbound

GEM: Generative Supervision Helps Embodied Intelligence cites this paper.

GEM: Generative Supervision Helps Embodied Intelligence Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:43:28.939856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:0f0d520a234c8e1c9750b03d4e4f4d68aa61b1fa34d144b327f0766844dc92d7

Observation 6657bccb-16d3-4b2f-b549-c6a75779ff99 · inbound

Wall-OSS-0.5 Technical Report cites this paper.

Wall-OSS-0.5 Technical Report Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:26:00.530454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T22:35:00.258436Z digest=sha256:57f3d16229ba190583fc466194f2712c272fa7dba1f96724a8d3b5099dbf24a9

Observation 6d129ccf-4039-44f3-a1d0-427ab9ba868a · inbound

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data cites this paper.

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.630648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T18:31:03.548002Z digest=sha256:56c20bd0c3ea5593bbaeb943c249945285ad97e151ced38f7bb078720c47ce04

Observation ea05806d-e451-4a73-87e4-ebb4d88fd28c · inbound

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation cites this paper.

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:34.770570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T06:26:32.209719Z digest=sha256:f395a3b57709c0ef5c505f9b2163730c00ca90482633e77d5309b690c9a07ecf

Observation 93a450d5-0507-48c7-8596-c619c4fb2a52 · inbound

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation cites this paper.

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T11:43:49.379947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:43:49.379947Z digest=sha256:4e95cbd67c80f8ad6b1f146e170d22d90596090caa521574493af7d3c4478ddb

Observation e9149a18-0ed3-4570-b74c-381e2a561269 · inbound

Vesta: A Generalist Embodied Reasoning Model cites this paper.

Vesta: A Generalist Embodied Reasoning Model Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:29:35.736977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T16:55:12.518255Z digest=sha256:ef9545328e09768e3053db9179a3eda9ddb7306b579350ec6dab507641bcc320

Observation 974cc6a3-9b1b-4118-884a-d95e485da3ce · inbound

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification cites this paper.

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:59:46.873792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T08:12:14.829556Z digest=sha256:bf59f0b3732ac806d14b085635d5c1b23ed911abbdef4005ae0cfeed416487cc

Observation a99f9c32-9729-45bf-81f1-806a6e7e4970 · inbound

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos cites this paper.

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-14T17:30:54.988498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:30:54.988498Z digest=sha256:c1bea46530688f8e4709bd6b8948952961f91e18103c3ae8cd1835cd6e9db327

Observation f7c34079-fe85-4677-91b1-0d98d1ed7f7e · inbound

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment cites this paper.

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T05:16:35.354226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:16:35.354226Z digest=sha256:13c31118ef1220e1ed543fb40817613d41b2a1d763e26a57c21dec5912c65e98

Observation 0ac0bafc-1318-471d-8f61-4b03a1718018 · inbound

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model cites this paper.

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T16:32:46.402572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:32:46.402572Z digest=sha256:26e315a33abca08358d46955bd4db3729474d107f0adcae6e549b1328d69c2e3

Observation 21826f6d-47ea-4d4c-be0f-410311188457 · inbound

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model cites this paper.

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T01:57:27.437916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:57:27.437916Z digest=sha256:ab23b036ff600b95f146028656dc3472665069b753f53721d76826650e96dd63

Observation 2b0413ca-47bc-4422-a25a-fed4ca82253d · inbound

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence cites this paper.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.179159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.179159Z digest=sha256:5037e4f6a067799de93b42169676c50cb6cedb84f02b283fde5021cfbbdba284