Pith. sign in

Paper Citation Record · LEDGER

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

As of 9 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 7 inbound Pith citation observations for arXiv:2502.09268.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09268 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:10:13.474149Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:26:40.586348Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:16:58.078690Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77857f45-ef83-4b28-9f0c-327c15996f64 · outbound

This paper cites MCIL trains a single goal-conditioned policy by mapping various contexts, such as target images, task IDs, and natural language, into a shared latent goal space.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation MCIL trains a single goal-conditioned policy by mapping various contexts, such as target images, task IDs, and natural language, into a shared latent goal space

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:13.954329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:10:13.439006Z digest=sha256:6dc7e566a9a4314b79e46a4238b0af605a5cc050fb1318171ce27ccf636f86d0

Observation 0907ae65-e66b-4c30-b0bd-77c67ab2996c · outbound

This paper cites MdetrLC uses text query mod- ulation to detect objects within images and demonstrates strong performance in tasks like visual question answering and phrase localization.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation MdetrLC uses text query mod- ulation to detect objects within images and demonstrates strong performance in tasks like visual question answering and phrase localization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:13.938766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:10:13.442947Z digest=sha256:3b79456a5d65a9d490ed5157e7b96ea905acb19f100c8bbb98679ae5bda37510

Observation 02ca56fb-f985-4f80-af86-2c27cf2bb354 · outbound

This paper cites an unresolved cited work.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T22:10:13.923892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:10:13.446944Z digest=sha256:4f95e1b1daddf70d6144e10dd647d830f07a299a2cf6b4a4259c6e5b47af36a7

Observation 0885ae0c-9d06-4c79-81c9-32b1fec08f36 · outbound

This paper cites Multi-column deep neural networks for image classification.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Multi-column deep neural networks for image classification

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.329891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.329891Z digest=sha256:d04b09cd739a7167c0fe240aa018d5b69f4b8fad7fc79e91f6a0825e6b5f2c85

Observation f1fd9f1e-c76a-437e-8e54-b594bbd20827 · outbound

This paper cites QUAR-VLA: Vision-Language-Action Model for Quadruped Robots.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation QUAR-VLA: Vision-Language-Action Model for Quadruped Robots

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T22:10:13.737757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:10:13.340871Z digest=sha256:609f59f679e3f50f0e3d336696a4637f741946bbdb626c5a858b45fcfe6b9784

Observation 2516f0e0-03cd-4a76-abb8-cbd4490c7864 · outbound

This paper cites Learning universal policies via text-guided video generation.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Learning universal policies via text-guided video generation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:14.015370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:10:13.346506Z digest=sha256:9970f27831b75530cf839ba2a8007d479c875ee854f4c8f4ee8c875a66a0ec5d

Observation 6fc987c4-dec2-4ff9-b155-f9ee0bbd2348 · outbound

This paper cites Denoising Diffusion Probabilistic Models.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Denoising Diffusion Probabilistic Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.356571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.356571Z digest=sha256:a851eba8f771cee3cc0d5fb21061d4455694085e0992d9395a18cac5cf642674

Observation b85776dd-69b6-4ec7-b660-b452a7b2d027 · outbound

This paper cites Learning to Act from Actionless Videos through Dense Correspondences.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Learning to Act from Actionless Videos through Dense Correspondences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.371215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.371215Z digest=sha256:b05ca382b30616591c538bf31a989ad714cb5cec109c70dc65f91bd44e84fa85

Observation 70c4fcec-1f1e-4baa-af95-4b606682321e · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.376171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.376171Z digest=sha256:9b2aaa9aff73fe41788cfa88d9f0abf48817fd4bde476de12f59870ad7ccd7d8

Observation c690418d-fc6f-48c7-9c61-2b72c2a73ad0 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.381137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.381137Z digest=sha256:f61b3b869238e57bd11f2aff17a5537d90a1e56b26f0b9cd07293774d5baf0ea

Observation 4cdcda9e-d353-4bc5-9939-bcd46c632d83 · outbound

This paper cites Language Conditioned Imitation Learning over Unstructured Data.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Language Conditioned Imitation Learning over Unstructured Data

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.386053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.386053Z digest=sha256:b2fc2b3d9de62e0827532ef05c444724de32523d2edd5338d297760631c2c3ee

Observation 7bfa824a-96b0-471d-9955-e41c25ebed75 · outbound

This paper cites What matters in language conditioned robotic imitation learning over unstructured data.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation What matters in language conditioned robotic imitation learning over unstructured data

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:14.000161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:10:13.391135Z digest=sha256:5fad719eec40bf818f3aece3a02d63c92129169551b9929fd1e523e48b4e1d20

Observation 3266e597-4108-4216-911f-0a1331347839 · outbound

This paper cites UL2: Unifying Language Learning Paradigms.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation UL2: Unifying Language Learning Paradigms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.406375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.406375Z digest=sha256:419f92a34a9af4566b67aad66229f406513198b36b54703f370b979c9acb8ce5

Observation 23772e90-b9b7-47b2-a94c-cbe4ca309cae · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.411250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.411250Z digest=sha256:45ac4b059451853a5ea40576c5e2bbf2fff100b8131ccda26a774aef8b3e33ce

Observation 0a303011-a4a4-4ae6-ac89-690b7b35470c · outbound

This paper cites Learning Interactive Real-World Simulators.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Learning Interactive Real-World Simulators

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.416313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.416313Z digest=sha256:b65083ca3ac4522f147a82ef24977665bc2fee84f0b5120322dad2936323855d

Observation ae8ed9be-3096-4d47-9f3a-a247369a29f0 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.421325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.421325Z digest=sha256:dd526f8f28b498c21e3564c05dd0782d842b590f9b3e3a5c778abb30ba75d50c

Observation 8fd42b4e-9e4f-410f-954d-d9f8695a8d1e · outbound

This paper cites Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.425586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.425586Z digest=sha256:f633f5154f783ae78911afee309214d907b755fa784614efc399d343ee3369f6

Observation a5b8a444-4ad4-40c5-a1e2-e671fe3d0028 · outbound

This paper cites RoboDreamer: Learning Compositional World Models for Robot Imagination.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation RoboDreamer: Learning Compositional World Models for Robot Imagination

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.429917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.429917Z digest=sha256:17b8e226864a632b8f6d748291fa71ab981261e7d0c81f7669111d209de6fe61

Observation f38028bd-bebc-4c38-bf91-1ff54971880d · outbound

This paper cites CALVIN Datasets.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation CALVIN Datasets

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:13.969688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:10:13.434172Z digest=sha256:f43ceb908d45e1d6aec76d0268f3dd111d52db73b4724c5a6d84b463c94165cb

Observation 916af6bd-da65-4567-99fb-987036e90c28 · outbound

This paper cites an unresolved cited work.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T22:10:13.908798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:10:13.451086Z digest=sha256:6e97bf08c12296cb7f4d5ca9cf7b31d3fda8ca9c0799420fb4d9cadad1757b3d

Observation 61136e6b-baca-498c-afb1-d98e0e5d757c · outbound

This paper cites This approach decomposes tasks into high-level planning and low-level action generation, improving task execution in com- plex scenarios.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation This approach decomposes tasks into high-level planning and low-level action generation, improving task execution in com- plex scenarios

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:13.892714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:10:13.455199Z digest=sha256:d1b5d371d933f10b1e065533590fa2ae5b1a47aa2a6c0f7ceac95bb62b5c91d3

Observation 7d8850be-c30c-42d4-838c-19eb376365b5 · outbound

This paper cites SuSIE integrates a large-scale internet visual corpus during sub-goal generation and achieves these generated sub-goals through a low-level goal-oriented policy.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation SuSIE integrates a large-scale internet visual corpus during sub-goal generation and achieves these generated sub-goals through a low-level goal-oriented policy

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:13.878128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:10:13.459443Z digest=sha256:09c707f550f4007123c7425adde3ea2eb418fe01e5c5918bd55bc3914a39a45f

Observation b1f26547-a910-458a-903e-73f8612dae7b · outbound

This paper cites time [s] 1 2 3 4 5 A vg.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation time [s] 1 2 3 4 5 A vg

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:13.862349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:10:13.464439Z digest=sha256:4c5dc01f662d2eaca2c4ba0275bb84dc52aa798232b755e0745e99e23415a59e

Observation 1447bb0c-990c-49be-a7c6-a4b93879d94f · outbound

This paper cites an unresolved cited work.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T22:10:13.831243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:10:13.474149Z digest=sha256:97b055f49cda27c54e8e66000181586efee2fcc4e5881b5b9b7ec2552d2ca9dd

Observation 79826b32-2b58-4ea7-bbf4-a72f253c24f5 · outbound

This paper cites an unresolved cited work.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Unresolved cited work

Reference 256

Resolution
unresolved
raw_fallback, observed 2026-08-07T22:10:13.846647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:10:13.469084Z digest=sha256:33e6ca08dc1206ab2628ee49f1a6f45c4ce8ad32a4660139073a2de64d39d09a

Observation f9a55c44-7a6c-4c07-846f-047ab736febd · outbound

This paper cites ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

Reference 1989

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.396381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.396381Z digest=sha256:ebe23217b6ff922adb00828166efb9c41ffb32fa9082020c9927efca7855c673

Observation 6fe32628-a306-4271-8b7e-b46bcac190a1 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.366340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.366340Z digest=sha256:aede26b610ccd634701223808faa468fd68f4018b614729a4d0e55909c89c2ab

Observation cbe8c123-ccae-4043-be10-6e7100ab296b · outbound

This paper cites Learn- ing to augment synthetic images for sim2real policy transfer.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Learn- ing to augment synthetic images for sim2real policy transfer

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:13.984953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T22:10:13.401709Z digest=sha256:af1c1c73c3ee9b37bdb95db49cac44532485b3fa1c4804626a309e1d41b97e55

Observation 1c6bc823-d6bb-4212-a499-aa72abd1804c · outbound

This paper cites High-Performance Neural Networks for Visual Object Classification.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation High-Performance Neural Networks for Visual Object Classification

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.335182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.335182Z digest=sha256:1c905cc8733629de250bf36dde3942c612e9d5dbe092fcb573484891f48cb7da

Observation c4e6c8f0-c41b-44da-a0b2-5ef64e776293 · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.324245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.324245Z digest=sha256:ebff24f8e825ebfb0bd23f7a9e2072b0a84c2c62248eae11490b8180a9383a09

Observation 096025b6-c5eb-4783-af4b-cc2115be1cba · outbound

This paper cites How Far is Video Generation from World Model: A Physical Law Perspective.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation How Far is Video Generation from World Model: A Physical Law Perspective

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.361758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.361758Z digest=sha256:3ecac84edaecf7cdb8a898908a0859cc54b0cff161f7d49925dca977867af64a

Observation 34b33d7c-b41a-42bc-8e51-54355e420024 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.306435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.306435Z digest=sha256:bfae7b2a0575df8f9ece308913a7ccf0cf0cc981a6618bcde116d2f9058b0255

Observation 392c608a-c55d-489c-9f06-83501e476f14 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.313181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.313181Z digest=sha256:8d4c8023c1657c277292044370f50b639b506a641cab4ef1f072a059e689f16e

Observation 1ea1bcbc-4357-45ea-923b-c6f37ff14cc0 · outbound

This paper cites Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.318398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.318398Z digest=sha256:84ac6cecd12ee07148853e86db43fb5b8042a1ee193ed847b7751969a4256194

Pith citing papers

Observation 087655ed-d631-4ecc-90a3-cb09436d4293 · inbound

Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation cites this paper.

Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:26:40.586348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:26:40.586348Z digest=sha256:101620b5980b5126d2e8d1ee88459019cf78fd5b62142c7a9f7ae39ec863c149

Observation a1b77edf-4751-4032-837f-11d91ee29833 · inbound

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning cites this paper.

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:25:59.007338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:12:08.970164Z digest=sha256:b83ee363d451fa906afd32440b9b37d659f9fa1c8d4611fcb7d3d50d2635977e

Observation dab481a5-713d-4ec9-b353-a4d72be0babf · inbound

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence cites this paper.

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-13T00:03:53.609175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:03:53.609175Z digest=sha256:b56223594f3e0dd7fc1fe20c41f3de9bb6e27369072602d529fcb1b099f63dc1

Observation e74777fd-9995-46df-84c4-9d70a46718ba · inbound

STRONG-VLA: Decoupled Robustness Learning for Vision-Language-Action Models under Multimodal Perturbations cites this paper.

STRONG-VLA: Decoupled Robustness Learning for Vision-Language-Action Models under Multimodal Perturbations GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:41:00.027103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:33:16.554905Z digest=sha256:159dc5afc717be24506c967dfe6ef329ab20e0daec1df747b2cdb09a89ad4f81

Observation 324a9723-1a61-4abb-b640-8aeace9bd246 · inbound

World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems cites this paper.

World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:30:19.071025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:25:57.218073Z digest=sha256:03ca37808c73df83a50bfa66726b54f50c82af0a786a2111d2779b939e3c1dc5

Observation 1693a345-199e-4622-a697-bec1e2ae1258 · inbound

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models cites this paper.

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:16:58.080059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T13:16:16.676244Z digest=sha256:c96f6c2adcfee752edb56ab0d06525d1ca8646e58f98890e9689d2209bef0a4e

Observation 8a465ad2-d513-480b-ab07-c2d1cc5d1253 · inbound

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation cites this paper.

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T14:32:58.543303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:32:58.543303Z digest=sha256:512b55916a9673d1cdf0dcd5627a7c1d8452cbb909313baa19c4b20590651dfe