Pith. sign in

Paper Citation Record · LEDGER

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors

As of 19 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2504.21530.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.21530 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:07:05.277615Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:57:26.612635Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:57:31.251745Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f074e8e6-36cd-4847-9400-191526bc441e · outbound

This paper cites GPT-4 Technical Report.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.014417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.014417Z digest=sha256:10dd025231a96f0ae2964ebcbd235b79e9c16ac6c3c45476297e88819426d3a2

Observation 0d461be6-308f-4049-b4c4-d4bf31ce7dd9 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors RT-H: Action Hierarchies Using Language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.020787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.020787Z digest=sha256:9ce816734f4a8007a949c7e51048d9b846c4d56ed2b730be47783f85fe72dc5f

Observation 0ef83400-73fb-493f-aebc-6736262fa895 · outbound

This paper cites Track2act: Predicting point tracks from internet videos enables diverse zero-shot robot manip- ulation, 2024.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Track2act: Predicting point tracks from internet videos enables diverse zero-shot robot manip- ulation, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:06.109718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.026019Z digest=sha256:3a63488fca9ed5c681e9c763b6da1781f559073f5d9b0044172ac68c3d8ff064

Observation 4836559d-3557-4069-804f-920f4084e771 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.030891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.030891Z digest=sha256:699d2f81f3cf010e9ef165537aeb6f29802e30a39a35bad031295bd94e2de647

Observation 8ce6648c-e311-409c-9a1d-44ef3561d056 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.036033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.036033Z digest=sha256:cd7335105324be48fee1e3feac1354cb884fbff11c1bdf8b16636362addfc5d9

Observation 2824a6d8-0541-4ffe-89e4-b040ff4cfe62 · outbound

This paper cites Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.041332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.041332Z digest=sha256:f75eb861f6bfc1324cb0f6623c8816e45431dae48d25945a946f5b6e104adc04

Observation 46e141c7-077d-4b04-9206-4332712dbd01 · outbound

This paper cites Shikra: Unleashing multimodal llm’s referential dialogue magic, 2023.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Shikra: Unleashing multimodal llm’s referential dialogue magic, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:06.093942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.046596Z digest=sha256:14fa83fb39445487b84fc8c007d4dff41b909c1fadb7e7ea4199fe5a8ecf8be2

Observation 69db985e-c967-4495-bffa-78c239e509c5 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action dif- fusion.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Diffusion policy: Visuomotor policy learning via action dif- fusion

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:06.078264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.051125Z digest=sha256:a32122f04a3bc6cd3b740bcfaf23c0f2806f440ec8e8d4e644488cc28d0d8ba2

Observation 61ba1ce3-9d6c-4b66-b236-ad5704926dfb · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.056513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.056513Z digest=sha256:ac391156ebee21114a0e84127921fd10f83c33d9a9b8a2bd82307a968f33a8a8

Observation cb84a066-c793-414a-9616-46a80439ccf0 · outbound

This paper cites Imitating Task and Motion Planning with Visuomotor Transformers.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Imitating Task and Motion Planning with Visuomotor Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.061342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.061342Z digest=sha256:977e93d57b9e682a2ccdb3e6516e54942ccc69cb3ae5c93b76d0aa11bc99fc84

Observation 5fb40c13-a3c4-425a-9999-85db878aed17 · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Objaverse: A universe of annotated 3d objects

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.066352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.066352Z digest=sha256:6c2d0c7908654126f266a1623ce77170e17c799e2fbd0bbb94ad68d4c7937b5a

Observation ee4c90f3-db52-4217-9d5d-df1c49c4340b · outbound

This paper cites FlowBot3D: Learning 3D Articulation Flow to Manipulate Articulated Objects.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors FlowBot3D: Learning 3D Articulation Flow to Manipulate Articulated Objects

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.071158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.071158Z digest=sha256:da84fde82e183f34e18b89caf0a00f43e59ad5440daf091a210aa534ccbdd5fd

Observation db5e5724-441e-4e18-a801-d1e2dfe57d50 · outbound

This paper cites Anygrasp: Robust and efficient grasp perception in spa- tial and temporal domains.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Anygrasp: Robust and efficient grasp perception in spa- tial and temporal domains

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.076200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.076200Z digest=sha256:2abc2f263275ac047f31d858b3cfb47b77de15a33ee7e2909aacffb2ef700721

Observation 67075254-6d1b-4dfb-b3a4-b71e2eac5973 · outbound

This paper cites Moka: Open-world robotic manipulation through mark- based visual prompting.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Moka: Open-world robotic manipulation through mark- based visual prompting

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:06.033158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.081370Z digest=sha256:487129f14b12b6158ddd1bb102592f847446029790417ffba7caa287d08618e2

Observation dce6ae1b-891a-4f92-acf5-519d726b0f15 · outbound

This paper cites RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.085786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.085786Z digest=sha256:1839081c0f0e82729e7d44732ddea7d1c6c2d81611d2c246fecd39535d4f92eb

Observation 1c440137-15a8-49d7-8e4e-808e75f09171 · outbound

This paper cites Masked autoencoders are scalable vision learners.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Masked autoencoders are scalable vision learners

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.090760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.090760Z digest=sha256:58e6341fe27399e2b9f9c2a1386ad0ff34f8d85a446eab1b5d09475c7d17e888

Observation 94e26d11-60ca-461b-9501-8ae73a4335c9 · outbound

This paper cites Perceiver: General perception with iterative attention.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Perceiver: General perception with iterative attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.095555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.095555Z digest=sha256:5a64b2f37a87e3a4e45a3b795953d7ad06e9d45255da3b9771a2441d22c25dba

Observation 2064a1ad-e8e5-4fd1-90f3-81836ef34b23 · outbound

This paper cites Rlbench: The robot learning benchmark & learning environment.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Rlbench: The robot learning benchmark & learning environment

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.998329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.100013Z digest=sha256:ed1292aa81e0e632c9e0c60b67b0255967242e1e5987d1d19d9ad86246394bc0

Observation 6dcf33ae-c757-4c84-9011-f240efd95410 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors OpenVLA: An Open-Source Vision-Language-Action Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.104695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.104695Z digest=sha256:940f497f44127388af49e46c662312679e17f98bd11bf5bdbeace21c28876946

Observation 2a032ea2-8a3c-4f81-9829-da414ccef191 · outbound

This paper cites Segment Anything.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Segment Anything

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.109396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.109396Z digest=sha256:9b0ace51938815a3f66cd010273089faf0c210886c1e2ebfba8448c9e9a8f07b

Observation cb65e6f4-5702-4e09-908b-9d2c06834f18 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Lisa: Reasoning segmentation via large language model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.983337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.114382Z digest=sha256:e9a92df2f333cfb866e29df82bd45ab3a92ebe90cf7682a12e183118de5a6358

Observation 8b745074-d37b-41b5-bb85-14f9fa64895e · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.118904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.118904Z digest=sha256:694c29ffa363e7c1905830c77529b1768ce5d3d4aee22ea57a5973572610a36d

Observation e331e974-d1ec-41bc-b1a0-a22077e7d56e · outbound

This paper cites Visual instruction tuning, 2023.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Visual instruction tuning, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.123999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.123999Z digest=sha256:6a9286c70cc7eba0ecd0d7cf8d10f094e10be7947bb367cc4b3ee6c584969fda

Observation e4e6f5ae-d174-4aec-aec3-fd7c21ea9676 · outbound

This paper cites Groma: Localized visual tokenization for grounding multimodal large language models.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Groma: Localized visual tokenization for grounding multimodal large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.948299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.128581Z digest=sha256:f20107771c9773432210ff02522796abf14df3288b3021e2bf4c70fd63eaafb6

Observation 03b09d70-72d5-49f7-8dcf-899593fbd439 · outbound

This paper cites What Matters in Learning from Offline Human Demonstrations for Robot Manipulation.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.133010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.133010Z digest=sha256:58e2de9723e332f1c3a474e6191257a18fd3adbe82beeae032bbda280c1cc444

Observation d7ad62f4-ff13-4623-9aaf-bdb8c8051f59 · outbound

This paper cites MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.137762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.137762Z digest=sha256:5296548be1073332d9d94e399764f30f620c2cd01fc597b41f30be03de45925e

Observation f5354ceb-51cd-440b-9a55-78753364ebba · outbound

This paper cites Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.143620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.143620Z digest=sha256:b0871cad48fd0d03e4435090f3928c7e68331c7165240a692c3bbb76a0e2521f

Observation 837e1179-462f-4f19-a917-b055c634894f · outbound

This paper cites RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.148191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.148191Z digest=sha256:15244063936f0ca270856081afeeb0222c44243b83614a9fd274c3d0bc0bbdd3

Observation c1715922-d1dd-4aab-9682-659c4505358a · outbound

This paper cites Kosmos-2: Ground- ing multimodal large language models to the world.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Kosmos-2: Ground- ing multimodal large language models to the world

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.922762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.153083Z digest=sha256:8989e3c51fb2e2fcdaff89149704a6ba59495bb2b97447fa380789d9d8233334

Observation ddd45b25-6d60-4b09-a78c-13707d575483 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Film: Visual reasoning with a general conditioning layer

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.907274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.157617Z digest=sha256:0a83cddb2cd8019ea5c2fc34f44a767decb6c8d0c71b29af2f23ecef3e34e2ab

Observation b2365aef-51a9-4f86-b86b-8a121f1917be · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Learning transferable visual models from natural language supervi- sion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.162217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.162217Z digest=sha256:eb3bbf3981708de1f84a8d95d17ec012ebee9932c4d042c82d29d8eca085f877

Observation 0eefa914-6846-4363-95fe-df26598a7180 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Glamm: Pixel grounding large multimodal model

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.881722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.166814Z digest=sha256:0261172ba51e05f509344f20e67f299ec1a4a9cd5ba525f5bfe6b0984dfc75b4

Observation 4775d94b-b1f0-4892-a54a-ec08ba8c47ff · outbound

This paper cites Toolflownet: Robotic manipulation with tools via predicting tool flow from point clouds.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Toolflownet: Robotic manipulation with tools via predicting tool flow from point clouds

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.171351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.171351Z digest=sha256:7a2a0b142ff9b2035a2607c14141d164a0a5f3a7e9afc4d5d4ed73d18f671d9c

Observation 6ea02a72-57ef-4be3-876c-634121302f77 · outbound

This paper cites Cliport: What and where pathways for robotic manipulation.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Cliport: What and where pathways for robotic manipulation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.856222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.175851Z digest=sha256:62574263f6de789b128fd6bda86451d56758e0d3b722e13277017e42baf9f281

Observation c4303925-de3e-4ff7-b20f-e2d8fda771b3 · outbound

This paper cites Open-World Object Manipulation using Pre-trained Vision-Language Models.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.180657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.180657Z digest=sha256:7548557dc69b96c0dbb76e334fe426999bf500d2c555c3ed7bfb02cc27d9aa2c

Observation 06adf63a-fb02-466e-917e-62a3edf1ff75 · outbound

This paper cites KITE: Keypoint-Conditioned Policies for Semantic Manipulation.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors KITE: Keypoint-Conditioned Policies for Semantic Manipulation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.185666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.185666Z digest=sha256:bd775b992adfa82316fa75cecc36d6348cfdd67db2ace1b8056f881cbd26c7e2

Observation 99ce5b5c-6f34-4606-8ad6-b484b9192006 · outbound

This paper cites Rt-sketch: Goal-conditioned imitation learning from hand-drawn sketches.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Rt-sketch: Goal-conditioned imitation learning from hand-drawn sketches

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.840640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.190494Z digest=sha256:b7fd8b8d2ffb2ebfc9a6205463d1c43f800da4111d7df74c349022a6b56dbb3f

Observation 3c951082-1103-41b6-a1fd-93714512b2d4 · outbound

This paper cites Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.195110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.195110Z digest=sha256:5bdf4ae9b57b8eff6c5677b63c95576cf5640df9d8ecb33c4aadea5f0b3d9e22

Observation aa50c0e3-b952-419d-8fb3-7bb6e17d6376 · outbound

This paper cites Robotap: Tracking arbitrary points for few-shot visual imitation.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Robotap: Tracking arbitrary points for few-shot visual imitation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.825296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.200214Z digest=sha256:e52c18dbec03a7dfbab0b3698ff1fefdda4d592c506e4788cf75f61a42611108

Observation 1f055d2a-0859-4feb-8b4d-d4d2e90921cb · outbound

This paper cites GenSim: Generating Robotic Simulation Tasks via Large Language Models.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors GenSim: Generating Robotic Simulation Tasks via Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.205184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.205184Z digest=sha256:3e8a47a0f68b297df99e163d7c311343294dec2235564f2bdd28e2874d631678

Observation c4c29b91-de0c-400b-b238-62856187c9ae · outbound

This paper cites The All-Seeing Project V2: Towards General Relation Comprehension of the Open World.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors The All-Seeing Project V2: Towards General Relation Comprehension of the Open World

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.209945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.209945Z digest=sha256:b4f021b682ac433776371bc94f5fe8722f02c4084b227899f73989cfce24db78

Observation 9c87fe91-406e-482f-9f33-dd3d4e553154 · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Any-point Trajectory Modeling for Policy Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.214926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.214926Z digest=sha256:c50d0fbf642e77271dec1aa6474badb20ab7c6759d6b3510419d4360377097af

Observation 38f1035b-a4ea-4dd2-91ed-9445a9c65e77 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.219984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.219984Z digest=sha256:837bd8f50f787f3e2ecf954e42b8ced4fc4cbce861dbcfa9688fc0f04d34f93a

Observation 7f07915e-3277-4d46-a68e-a438cba7f124 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.809925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.224660Z digest=sha256:f9a5cb2e05ce43f1d0935ef4abcdd0492f821ade44ba00c24411b8e065fe437f

Observation 029ea9f6-7685-4cf3-9bce-2e865967315b · outbound

This paper cites Flow as the cross-domain manipulation interface.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Flow as the cross-domain manipulation interface

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.794214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.229333Z digest=sha256:cae7c5ed04cd2b220699633e94d7352cfde7a018dbec45c989895c33627cc65f

Observation ecbbc05a-5278-44fc-8150-7e04daaf905a · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.233862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.233862Z digest=sha256:d6a79389ef66b84db8fb7520a31a83dd9bcfafd6eadacf746ddd2b432a7b7a7b

Observation e7e61a4b-e04a-4e2f-8d5c-e7aee350d623 · outbound

This paper cites General Flow as Foundation Affordance for Scalable Robot Learning.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors General Flow as Foundation Affordance for Scalable Robot Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.238702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.238702Z digest=sha256:791bf3750ec3f9c3eb9aa45e31714e160a62550056ef430d7b61d98d3e6b44d3

Observation 169caaa6-595d-4751-b87d-a9bd1b76cd8c · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.244349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.244349Z digest=sha256:e7ec08f9a6abc8ec3f324bcb7bddae6d528c2824fd8cce03719330db57b6d71c

Observation 7c18a2de-149d-4851-8d80-6174614df944 · outbound

This paper cites Sprint: Scalable policy pre-training via language instruc- 10 tion relabeling.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Sprint: Scalable policy pre-training via language instruc- 10 tion relabeling

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.776430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.249417Z digest=sha256:d81a44bbfae86f9df3f59bc625f3bea28591d762e31beb361d63383dacd54d4a

Observation 6872fe4e-93ad-48de-b552-228b094bfe3d · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.254066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.254066Z digest=sha256:3e15f35ddf2719c3d930a5b06e2ba7dbbad309df188ec3c4094f7b77a6e24573

Observation 9ae6efb6-9784-4b30-86a4-c8cf4450b3b9 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.258931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.258931Z digest=sha256:99921523cc666c152127c4b7d2387105aa4c38ac5b2a9262fd3ccf13d4ce7b9e

Observation 4c314359-1979-466b-9693-1c318cd8b848 · outbound

This paper cites an unresolved cited work.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:07:05.760836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.263792Z digest=sha256:d63719222b946ec5b33884cb72e0983692a38d94a53f3b1b48207456f6f1e603

Observation c5833f51-b78b-47a2-9536-ba3c34ab632e · outbound

This paper cites an unresolved cited work.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:07:05.745175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.268368Z digest=sha256:fd822773f7fb401be7fcce6ba08e64d293a2121b659f4d26fdab7256979128fc

Observation 71b8116e-2fb6-44ad-bc13-402a9bbaa1d2 · outbound

This paper cites an unresolved cited work.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:07:05.729786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.273044Z digest=sha256:86bd1cef65d32becbf314cdcaaa22e84d1564856b01d2b42e5e8d642c0b4575f

Observation a133c291-10b0-4dfa-b7c8-e906210a3e78 · outbound

This paper cites yellow surface with brown spots.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors yellow surface with brown spots

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.714651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:07:05.277615Z digest=sha256:1ab0df0f5f05f9005ab5099fb24fced9a2ea33e3fd5f8fdf8c005f7a27485664

Pith citing papers

Observation 7c29c10a-ade1-4c72-91e6-45a7d04c8a88 · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RoboGround: Robotic Manipulation with Grounded Vision-Language Priors

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:57:31.256067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:57:26.612635Z digest=sha256:219472f66f0a2a347b2f7e0172236e531ab42224f0aaf93ad1c4df361103e584