Pith. sign in

Paper Citation Record · LEDGER

ROSA: Harnessing Robot States for Vision-Language and Action Alignment

As of 16 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 2 inbound Pith citation observations for arXiv:2506.13679.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13679 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:31:31.854787Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T01:01:44.128884Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-08T01:01:44.731686Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact3
  • verified fuzzy12
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c814ad01-a6d5-4cb3-92ba-6f0b2c45367a · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.706918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.706918Z digest=sha256:c5e44f69c4037fcef457b72158d4313ee15dc61af00f69b3cab991bcefe627e8

Observation b09b29b0-0459-461f-96db-ea6356a82b61 · outbound

This paper cites Improved baselines with visual instruction tuning.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Improved baselines with visual instruction tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.710794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.710794Z digest=sha256:9ba931eea00b449b801d45f3838d7c795364ec00cf72ffb8e98db27f4c335be7

Observation 2997546c-b778-430d-bf0c-43a9f82d0535 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment LLaVA-OneVision: Easy Visual Task Transfer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.714299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.714299Z digest=sha256:45d735d1192c6d05dff1155b39919cb1eb80a8c664f77fcb006a0cbbaef00b84

Observation 904717a8-69be-4bdd-a694-e1085e59ddd0 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.718384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.718384Z digest=sha256:544407db699e44ba23a8d818d19f4577007d6b441dfc7c8a11bb8370791563e1

Observation c58b818e-dfaf-4cd2-a62e-d3fdd6d3fdc1 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.722611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.722611Z digest=sha256:79c94a532ed917f2178beda982b0f329e5abb0f493f73f17ff475545d0d8a98d

Observation bc828413-b729-4b96-b211-90854633e3c7 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.727134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.727134Z digest=sha256:0674137a298c94ce6271cb4b0eb385dd298ae45d586fcc385fe3785d3e6e21b4

Observation d8407032-91d5-48bb-a752-8ba9c45b3204 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment OpenVLA: An Open-Source Vision-Language-Action Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.730726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.730726Z digest=sha256:c91af2703daf57f736379f43c0b3e6b2c1ea196c409aa1edeff941749adc6927

Observation 2fbb4476-14fa-406a-b8d6-a6499681c0fa · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Open x-embodiment: Robotic learning datasets and rt-x models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:32.187358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T00:31:31.733936Z digest=sha256:c169048d7012ec63826a994d76e161614721cddca34e74d9a7538d28d42058ba

Observation 3c4898bc-9a6a-4a92-9291-2be204bd8d2f · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.736995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.736995Z digest=sha256:90497e992a15f04099b4beb811be471b4f05fb949038c2a0ba34f7c34eb6a6a2

Observation 6b56d11b-00c3-4ad2-aa8b-4ced69a563dd · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.740812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.740812Z digest=sha256:ac84cc826bdd2360ef618b9b7a9927e9165f820df2a655f82e4a40eb2f1a8e43

Observation 165a2863-afd9-46ef-9a11-eeb195aa3bcd · outbound

This paper cites ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.745215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.745215Z digest=sha256:f87af1340a9788ff630c5f94cc66f14e96c77e92155c0b02391948b2a1f9ffae

Observation 0fac4feb-cba5-4cc3-b41a-23c6ed007222 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment RT-H: Action Hierarchies Using Language

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.748910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.748910Z digest=sha256:2d7541bfa8c3233fb57b86f55cd1af816a13968c5131327a41352d98e8ff2c4c

Observation 74dab451-b4d4-4257-b8c1-0d5ecffc4965 · outbound

This paper cites Octo: An open-source generalist robot policy.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Octo: An open-source generalist robot policy

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:32.179617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T00:31:31.752768Z digest=sha256:3ba4442d048342e571f04e630e1ff71fa584905d41c1c787d21c030c0a717192

Observation 59b08fe7-ffbe-492e-9fcc-56f8c2beba24 · outbound

This paper cites Perceiver-actor: A multi-task transformer for robotic manipulation.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Perceiver-actor: A multi-task transformer for robotic manipulation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.755909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.755909Z digest=sha256:f9095ee21ce5d7b77d826c92b5e3ca28bac31916df779185f7c372beff675306

Observation 7b4db716-9f2c-4606-9d99-6282f0170e54 · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Vision-Language Foundation Models as Effective Robot Imitators

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.758456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.758456Z digest=sha256:88d67b92c0736efe2429a7a78c725e24e8b02b05f0556c3d48c7193723debb37

Observation b992ffc5-e23c-4980-81b7-5aa8b2145abe · outbound

This paper cites Rvt: Robotic view transformer for 3d object manipulation.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Rvt: Robotic view transformer for 3d object manipulation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.762158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.762158Z digest=sha256:afc8a0da49f9ba7a604038389f9d23180f59ed2d118d4340a42005d17693c39c

Observation d15c9dcd-c39e-4e06-8e8d-338192f51988 · outbound

This paper cites RVT-2: Learning Precise Manipulation from Few Demonstrations.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment RVT-2: Learning Precise Manipulation from Few Demonstrations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.765815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.765815Z digest=sha256:ea1aa31b85ff63dff0bad6d7c293234105c1e006cf75af37367202ad85ed8d11

Observation 3bfe38de-0259-45f7-b169-78caaec5f641 · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Bc-z: Zero-shot task generalization with robotic imitation learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.768857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.768857Z digest=sha256:25f584db568226872a17c6600e8dbaa9d23ebc893ebd929cb26f9401673d9ced

Observation 0b281560-27ce-4836-b863-aace77d3cfc0 · outbound

This paper cites Coarse-to-fine q- attention: Efficient learning for visual robotic manipulation via discretisation.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Coarse-to-fine q- attention: Efficient learning for visual robotic manipulation via discretisation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:32.157498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T00:31:31.771680Z digest=sha256:826aac533c9cdb7369a99aa1e1e6dedc808e8971440dfbed62ccdf16c97d3a83

Observation e2a4cdef-d831-43af-b315-087c08719fd5 · outbound

This paper cites Robot learning with sensorimotor pre-training.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Robot learning with sensorimotor pre-training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.774404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.774404Z digest=sha256:55cd53b09bd26132b03492ca2c06244efb915209eb8f4bc796d272c5339c93bd

Observation 0a70a004-3e9b-46d0-a405-26dc6d678029 · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment VIMA: General Robot Manipulation with Multimodal Prompts

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.777184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.777184Z digest=sha256:ee93b92b2b10df9d0db709c0c5a717f98e1079d02ec0aec0c8b09efc153693c3

Observation fd92e520-1295-4fd6-aee2-c3b5429fdcff · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment RT-1: Robotics Transformer for Real-World Control at Scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.781063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.781063Z digest=sha256:9433efa39d7253f4a2a068da14290daf29fb9e483cf7426e9c5e4ff1c58884d5

Observation 1ae0e367-31f4-4e7b-b225-0d48e47ac061 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.784315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.784315Z digest=sha256:bc3d3e83a2a5eacbede2eb750c461cb63143781430c871ef84efe032eeb73bf8

Observation acb22942-725d-4b03-9e9d-ad6442942450 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment PaLM-E: An Embodied Multimodal Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.787616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.787616Z digest=sha256:e88be1297fd89368451dd77b31d8a211945255739af17e4449eabce25b06f47f

Observation f1192e56-1c1f-429e-8e4e-dcdfc069c8d9 · outbound

This paper cites Learning transferable visual models from natural language supervision.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Learning transferable visual models from natural language supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.791156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.791156Z digest=sha256:fabc522a053a8cc44eff4d8c1cad88bc8fa1fc146275dc183e8c5e7c875e1fc4

Observation 8589d5cb-f1e4-4b80-b6c2-c6a8b42f89f1 · outbound

This paper cites Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.794838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.794838Z digest=sha256:ea2e65e11685f68c3d03bd727d3a0d8f691ab5a6d99722b3eb859527d1f94dc0

Observation ef2fa0cb-c995-4d5e-8dc1-51545a5300d9 · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.797767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.797767Z digest=sha256:23ec5b6defcd1e2d29e7173858baaafed7898430604a63ed00f52e8b2b72eb2a

Observation 3f6cc393-fa1c-4e30-bbae-f03991b8191f · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.800648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.800648Z digest=sha256:02e3adead1b0496942521fe6b06dc9d86dca203bca52f3b4dc9d98c1a321b965

Observation 41f07b63-3605-43c8-a2ac-81775e33644d · outbound

This paper cites Llarva: Vision-action instruction tuning enhances robot learning.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Llarva: Vision-action instruction tuning enhances robot learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:32.135298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T00:31:31.804344Z digest=sha256:4673166bd8c2e6d271e4171021e03d70863ca57054511b13e55c6bfcace4734c

Observation dd50b737-8aac-407b-a9c3-55f3d821b0cf · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.807156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.807156Z digest=sha256:40a451754107f7d1461b87f1d4f801e585e8f93a8eba353ed2a462989aa1f8ac

Observation a53aae7b-3c62-416e-9a3a-fc5a03f3ecc1 · outbound

This paper cites LLaRA: Supercharging Robot Learning Data for Vision-Language Policy.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment LLaRA: Supercharging Robot Learning Data for Vision-Language Policy

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.810526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.810526Z digest=sha256:957b38306bb90a530a1b4ca55619c93fa51f47aad27aad820b37460fd33a4ca1

Observation ced5fe00-6396-4a32-8e96-f20f57f39cb0 · outbound

This paper cites Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.IEEE Robotics and Automation Letters, 2025.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.IEEE Robotics and Automation Letters, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.813608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.813608Z digest=sha256:4ee795fa8c1b8b9fe964cb837582235dc5d6ada896756400a3bcdc4f57025307

Observation 1090907d-ac6e-4a27-ad05-3b22f1083705 · outbound

This paper cites Towards Fast, Memory-based and Data-Efficient Vision-Language Policy.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Towards Fast, Memory-based and Data-Efficient Vision-Language Policy

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:31:31.931357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T00:31:31.816613Z digest=sha256:da356d860cb6c0689bc850ba5716c162b84b01cb45ef95ba781d0fb4f340382b

Observation fbe82f9c-440f-4a16-9397-b85d4f3eb02a · outbound

This paper cites MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.820504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.820504Z digest=sha256:22fc9b4a25eb290b34b09d2f87e71816a4f0224abe4607e12d27884eb44b46e2

Observation ee6b8ece-75f1-4d92-8a42-ee1ae762e9fc · outbound

This paper cites An Atomic Skill Library Construction Method for Data-Efficient Embodied Manipulation.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment An Atomic Skill Library Construction Method for Data-Efficient Embodied Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.824023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.824023Z digest=sha256:32b27f1fc82c7d4a2153ad306fea399b6c4c3f850f8a8914b7907fd2a6a1a0df

Observation b74616f8-37a9-4c9f-91d3-09548dfd9de0 · outbound

This paper cites Pose estimation for an autonomous vehicle using monocular vision.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Pose estimation for an autonomous vehicle using monocular vision

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:32.121947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T00:31:31.827326Z digest=sha256:160875aa5262b9769e487b1955b6529a1305d3f2368769977fa70f48791b993a

Observation 0e7fec20-bfcd-423a-8a42-f529c0193c35 · outbound

This paper cites Pose estimation and map building with a time-of-flight-camera for robot navigation.International Journal of Intelligent Systems Technologies and Applications, 5(3-4):355–364, 2008.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Pose estimation and map building with a time-of-flight-camera for robot navigation.International Journal of Intelligent Systems Technologies and Applications, 5(3-4):355–364, 2008

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:32.113138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T00:31:31.829691Z digest=sha256:faa4f1bdd2e57ec6dde25c0f2e523a371e5ff02282437e5463be840c5f5fa15b

Observation 0a41b70d-1cae-4cc4-9308-da39c0a7c1d1 · outbound

This paper cites Learning human-to-robot handovers from point clouds.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Learning human-to-robot handovers from point clouds

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:32.104564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T00:31:31.832008Z digest=sha256:26cee87e0f4a97f279aa175db9428d76fd94023578490817b8ef15c5e620e7d6

Observation a55ac823-aecf-4896-b0f8-5f943267b7d9 · outbound

This paper cites Pose estima- tion and adaptive robot behaviour for human-robot interaction.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Pose estima- tion and adaptive robot behaviour for human-robot interaction

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:32.095313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T00:31:31.834485Z digest=sha256:63ed9b1d7ae3074bf61e9c158093b121f950658a8f46585749e9e78ba6cc3ad7

Observation 7401f7d3-8dbf-40dc-9761-86f2867a01d5 · outbound

This paper cites Privacy-Preserving Pose Estimation for Human-Robot Interaction.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Privacy-Preserving Pose Estimation for Human-Robot Interaction

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:31:31.906391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T00:31:31.836905Z digest=sha256:b72d4ae470c9653cec8f295123719c92ee8c5eeb3cbd4167d5efde57b68a49db

Observation ec4159cd-4cc7-4493-9087-514d9950cc0e · outbound

This paper cites 3D Robot Pose Estimation from 2D Images.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment 3D Robot Pose Estimation from 2D Images

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:31:31.894679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T00:31:31.839905Z digest=sha256:f395a3a12d8b977750f311d3e38135e5708e0bfcccdbc41bb999a25676911cdb

Observation 24dd54f8-a3d3-4790-a295-2b17553f9b5d · outbound

This paper cites Multi-objective convolutional neural networks for robot localisation and 3d position estimation in 2d camera images.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Multi-objective convolutional neural networks for robot localisation and 3d position estimation in 2d camera images

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:32.086060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T00:31:31.843312Z digest=sha256:53cac7b5369c4ded687e62032d253579e4c400e60aa40b25dcb6e84ad638c109

Observation 688cad5f-3aba-4f4b-956b-ee7eb8eda90f · outbound

This paper cites Robots’ state estimation and observability analysis based on statistical motion models.IEEE Transactions on Control Systems Technology, 30(5):2030–2045, 2022.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Robots’ state estimation and observability analysis based on statistical motion models.IEEE Transactions on Control Systems Technology, 30(5):2030–2045, 2022

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:32.076280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T00:31:31.846316Z digest=sha256:0f61125e04dd72e903d5bf1a3de13c6170a399d9117bda09ffd3de4088b36af3

Observation bc44f804-0131-4409-8eac-bd62b734ec53 · outbound

This paper cites Real-time holistic robot pose estimation with unknown states.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Real-time holistic robot pose estimation with unknown states

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:32.067024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T00:31:31.849212Z digest=sha256:81e407ab02877a6da9bd9d24af638b85c3494d33b329963d211e2c1390342da0

Observation 65460fd1-5104-4768-bee8-b9ebbc0aa0ab · outbound

This paper cites Qwen2.5 Technical Report.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Qwen2.5 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.852169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.852169Z digest=sha256:bab633e9a78674546f254fa54fbf6adba125e02e38c9ef78bc950b97987d1d96

Observation 534a41be-be97-4cec-aa30-2727b67db404 · outbound

This paper cites push the maroon button, then push the green button.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment push the maroon button, then push the green button

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:32.057279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T00:31:31.854787Z digest=sha256:c4868586cf009ceca33ddf43d5fac58962be6a4a332281cf34028c29d7ee7c03

Pith citing papers

Observation 806cce81-a07d-4a0d-964b-55c4be1e767f · inbound

Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models cites this paper.

Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models ROSA: Harnessing Robot States for Vision-Language and Action Alignment

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T06:10:35.151686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:10:35.151686Z digest=sha256:13c27621e952cfcf73e8a50665bf97958cbbbb40b66f8a7073ac67648d5a382f

Observation 87fdd124-616f-4096-9bdc-3ba1d7801856 · inbound

How Should Vision-Language-Action Models Use Proprioceptive State? cites this paper.

How Should Vision-Language-Action Models Use Proprioceptive State? ROSA: Harnessing Robot States for Vision-Language and Action Alignment

Reference 2020

Resolution
verified exact
local_arxiv, observed 2026-08-08T01:01:44.735748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T01:01:44.128884Z digest=sha256:e0900c493085944da0ef9c2e03e0bb6e71a1f416c34eb39a1c18aad714615bd5