Pith. sign in

Paper Citation Record · LEDGER

Training-free Generation of Temporally Consistent Rewards from VLMs

As of 9 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2507.04789.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04789 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:47:44.014826Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f134ada1-780f-48a4-be4d-1f63cc164808 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Training-free Generation of Temporally Consistent Rewards from VLMs Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:40.030491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:40.030491Z digest=sha256:94b7a28055f9d2e44933431dc091d8d4d08243b2c5c45e6caf01471ff15af98b

Observation a287c5c4-e6d8-4dfd-8dfd-0407e8d7bb64 · outbound

This paper cites Vision-language models as a source of rewards.

Training-free Generation of Temporally Consistent Rewards from VLMs Vision-language models as a source of rewards

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.972386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:40.080984Z digest=sha256:ce964019867f296a675fe8421df27c4c7d6779d501757ca3eff190656aff9d71

Observation 4b02908f-fac7-4629-99ba-9633105e81b9 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

Training-free Generation of Temporally Consistent Rewards from VLMs Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:40.142313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:40.142313Z digest=sha256:a167df8d7fecf3a459e0aa4c5496767d020377fb43a1b67e63a7f2b694b90775

Observation f1d4c3a6-72d9-47a9-be9a-8ed80b7cc642 · outbound

This paper cites Towards a unified agent with foundation models.

Training-free Generation of Temporally Consistent Rewards from VLMs Towards a unified agent with foundation models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.809780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:40.179746Z digest=sha256:7d6d42e0f56bfde690abd0203d9378ab7d11f7cdf8ddf5f11506b6e588a2807f

Observation 23a79ce5-2d5d-46e7-8591-be13e497db7e · outbound

This paper cites Video language planning.

Training-free Generation of Temporally Consistent Rewards from VLMs Video language planning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.651023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:40.263441Z digest=sha256:e615dce690c61c7fab895bf9b5ee3ee670f6dcba88f7a11644326d0bfaa5f8c1

Observation a4937b0b-b98f-47d0-b8da-e5b9fa2e992d · outbound

This paper cites Manipulate- anything: Automating real-world robots using vision- language models.

Training-free Generation of Temporally Consistent Rewards from VLMs Manipulate- anything: Automating real-world robots using vision- language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.455521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:40.331962Z digest=sha256:ac1858cb3fe3df8582e0d2aafbb3c7fd1cefb7a9810502c7b4f971b74e266c8d

Observation f0f9be62-7248-4a83-b51e-bff7ba31dddf · outbound

This paper cites Phys- ically grounded vision-language models for robotic manip- ulation.

Training-free Generation of Temporally Consistent Rewards from VLMs Phys- ically grounded vision-language models for robotic manip- ulation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.267732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:40.421741Z digest=sha256:6f41d95af8e3af03bfcb2740182055ba213a3b0321d16a92673fccc8214c20d1

Observation 46fd493f-6084-4191-90ea-d1646bb4cc50 · outbound

This paper cites Doremi: Grounding language model by detecting and recov- ering from plan-execution misalignment.

Training-free Generation of Temporally Consistent Rewards from VLMs Doremi: Grounding language model by detecting and recov- ering from plan-execution misalignment

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:51.078631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:40.507364Z digest=sha256:fb53f6bfbb889be6fda9525d01b27cc081581a55a7ba4af52479a6cfdab4e57a

Observation 909f63de-b7b0-47fa-86d9-2c31311c91f9 · outbound

This paper cites Mixgen: A new multi- modal data augmentation.

Training-free Generation of Temporally Consistent Rewards from VLMs Mixgen: A new multi- modal data augmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:40.622191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:40.622191Z digest=sha256:3d09fafabe6e270227a74cd05abb2064b65b766ae633c09b06cbad2824bfbcac

Observation a69a73d2-41fb-4632-add1-e48336247fac · outbound

This paper cites V oxposer: Composable 3d value maps for robotic manipulation with language models.

Training-free Generation of Temporally Consistent Rewards from VLMs V oxposer: Composable 3d value maps for robotic manipulation with language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:50.920534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:40.668131Z digest=sha256:5d4780104cd39debb208a3e000ba244e246a3b4bba9800b6a0435f7e6e1c665b

Observation a2f238db-8987-4812-911c-21b459dfa134 · outbound

This paper cites Robobrain: A unified brain model for robotic manipulation from abstract to concrete.

Training-free Generation of Temporally Consistent Rewards from VLMs Robobrain: A unified brain model for robotic manipulation from abstract to concrete

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:50.753133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:40.739804Z digest=sha256:d16772254a218a5572c51ef0915db6cc0f66a906300af0af3e8110123ade4eee

Observation 1736832a-c8e6-4a07-bbb1-29512151ca3f · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Training-free Generation of Temporally Consistent Rewards from VLMs Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:50.534051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:40.807513Z digest=sha256:32a79c3df5a153ff366f5b101b92387d4bddaed216039ccea0ac427fb1aff81c

Observation 6bbb4804-88f2-4259-bfe3-7ac2b93ef29f · outbound

This paper cites De- composed prompting: A modular approach for solving com- plex tasks.

Training-free Generation of Temporally Consistent Rewards from VLMs De- composed prompting: A modular approach for solving com- plex tasks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:50.298515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:40.867070Z digest=sha256:be50113520dffaab49fc26d76205cfe233ed5a86d4fa48288066b04155e0b261

Observation 768e7873-9fd7-4bfe-9814-3b021123782b · outbound

This paper cites Segment any- thing.

Training-free Generation of Temporally Consistent Rewards from VLMs Segment any- thing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:40.928757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:40.928757Z digest=sha256:61d25083fb96ece1be962dc8bafe5985a61673714cf17261688073995268e5b3

Observation 3360e7b3-df2e-4695-8faf-1822753c8fbe · outbound

This paper cites Multimodal sensor fusion with differentiable filters.

Training-free Generation of Temporally Consistent Rewards from VLMs Multimodal sensor fusion with differentiable filters

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:50.117560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:40.986288Z digest=sha256:d3d1d56f20ca3475c87b756ce7648a6065e07a6327c43d67d6fc2ff51067bd06

Observation 74fe2b5e-accf-4561-8206-e9516c154ca0 · outbound

This paper cites What foundation models can bring for robot learning in manipulation: A survey.

Training-free Generation of Temporally Consistent Rewards from VLMs What foundation models can bring for robot learning in manipulation: A survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:41.045499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:41.045499Z digest=sha256:9d2332a7520e4a01bdd88e3fa328d3e09686b4e2a0d485af37867943eca00199

Observation f6987d0b-d893-4fc4-86ff-5bfd90c13d27 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Training-free Generation of Temporally Consistent Rewards from VLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:49.856803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:41.100895Z digest=sha256:ae51a31094d2c6cfd00bcc5f0944fd55488b646850b33369f9157c67808e94c1

Observation 469d33dc-beb8-457e-b52c-98824cd12f77 · outbound

This paper cites Code as policies: Language model programs for embodied con- trol.

Training-free Generation of Temporally Consistent Rewards from VLMs Code as policies: Language model programs for embodied con- trol

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:49.716556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:41.135694Z digest=sha256:bf067cd8d1b87c8799d4318fcf0dcd0bbdbe31c79c3a0bef7a2c38adc1aca69a

Observation be3dc320-6d76-4df1-8112-9b1642ac2f8d · outbound

This paper cites Reflect: Summa- rizing robot experiences for failure explanation and correc- tion.

Training-free Generation of Temporally Consistent Rewards from VLMs Reflect: Summa- rizing robot experiences for failure explanation and correc- tion

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:49.579110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:41.187899Z digest=sha256:5bdff9f93fd7759cce89c47c8f32d6997e4a6a8e717b20f0cfd1d264079aabce

Observation 41423878-e643-4d71-a570-f1945e5c1435 · outbound

This paper cites ExploRLLM: Guiding Exploration in Reinforcement Learning with Large Language Models.

Training-free Generation of Temporally Consistent Rewards from VLMs ExploRLLM: Guiding Exploration in Reinforcement Learning with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:41.236915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:41.236915Z digest=sha256:7b39f875c59ecc2573de9f813973adff753e4b7764eaa5b575832a40fde9d6e2

Observation a3bcff94-6a78-43d9-9dba-0aea345740ba · outbound

This paper cites Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks.

Training-free Generation of Temporally Consistent Rewards from VLMs Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:49.277646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:41.292412Z digest=sha256:240a8cacf997e3c67a7d0d91be318856f17b4bbd6c9eadd639c5c40997638c16

Observation baaa893d-4bfb-4997-956a-a2b1f8e47150 · outbound

This paper cites The deep latent space particle filter for real-time data assimilation with uncertainty quantification.

Training-free Generation of Temporally Consistent Rewards from VLMs The deep latent space particle filter for real-time data assimilation with uncertainty quantification

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:49.089889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:41.337357Z digest=sha256:19e433650a5668b7366d83582a83efc37fd45fac0b4de3dbed5af1dddc2405de

Observation bb729636-72e8-4237-8032-cbae885bfbfc · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:48.903738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:41.384122Z digest=sha256:43d12afbddac90e6c848720ce3fa158d810dec0c436897acde3ab3b0c200acce

Observation ef9ca880-c2e3-4d85-9b2c-7aed41380125 · outbound

This paper cites A real-to-sim-to-real approach to robotic manip- ulation with VLM-generated iterative keypoint rewards.

Training-free Generation of Temporally Consistent Rewards from VLMs A real-to-sim-to-real approach to robotic manip- ulation with VLM-generated iterative keypoint rewards

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:48.721460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:41.447612Z digest=sha256:6cddb675b857fdc5c44e3a16e3313f8bf8c8c4879ad3eaa7811b8a259d29d196

Observation 99187c9c-c470-45b2-b0aa-b6df514ad934 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Training-free Generation of Temporally Consistent Rewards from VLMs Learn- ing transferable visual models from natural language super- vision

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:48.543814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:41.501344Z digest=sha256:626de658284c0e893cc60911a256459643a753fd18b61c9bcff21fb758a42ca3

Observation f1662e6a-7c88-47f7-b24c-85f183ef3f94 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Training-free Generation of Temporally Consistent Rewards from VLMs SAM 2: Segment Anything in Images and Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:41.565323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:41.565323Z digest=sha256:c23b8bc5b6c50c11ca209e826efe74f3e2b4ea31bf24d62bc683642c3db6e7ac

Observation 23706177-03b2-4fbf-8698-4cba3fcd5a9d · outbound

This paper cites Vision-language models are zero- shot reward models for reinforcement learning.

Training-free Generation of Temporally Consistent Rewards from VLMs Vision-language models are zero- shot reward models for reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:48.405348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:41.604704Z digest=sha256:3d51e5b8efa5ea0e5f0b835092d66c6628b8a76ad2f5f47e2772205968cc0564

Observation b25714cc-0c6e-47d5-8b07-a3043dcb36dd · outbound

This paper cites Chinchali.

Training-free Generation of Temporally Consistent Rewards from VLMs Chinchali

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:48.164013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:41.664896Z digest=sha256:b611c1e9f85a15b8b5edfe6273ed45f469d53a1499ba6a8a5ba7cf501b3f819a

Observation 10d0986b-78cb-4f54-971c-ac9ea122b65e · outbound

This paper cites Cliport: What and where pathways for robotic manipulation.

Training-free Generation of Temporally Consistent Rewards from VLMs Cliport: What and where pathways for robotic manipulation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:47.942743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:41.728121Z digest=sha256:7c56fba895472a40ac68b7cef88f3e462624b9ead6330df524a816d0ad237a0f

Observation b1f21a69-874f-41b1-b189-c8ff53451e41 · outbound

This paper cites Perceiver- actor: A multi-task transformer for robotic manipulation.

Training-free Generation of Temporally Consistent Rewards from VLMs Perceiver- actor: A multi-task transformer for robotic manipulation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:47.735982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:41.817923Z digest=sha256:b17f4cb5c6c40561879b3cadc56be224f73e9f74d74c2b9e0795ca7d677d0750

Observation b4c1f25f-8f94-4208-8291-1bdc5dcf0673 · outbound

This paper cites Progprompt: Generating situ- ated robot task plans using large language models.

Training-free Generation of Temporally Consistent Rewards from VLMs Progprompt: Generating situ- ated robot task plans using large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:47.543615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:41.891782Z digest=sha256:7cc48e95f239594d96eb7df2bd632e233d21fa7027b1a7a303719bb0abba12ce

Observation c54b97b5-9066-497b-999e-157c0dd7a806 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.

Training-free Generation of Temporally Consistent Rewards from VLMs Reason-rft: Reinforcement fine-tuning for visual reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:41.937872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:41.937872Z digest=sha256:1f6e1a3fe9fbeae7484f997efdf7058b12e039d2213483013c84cf9d85891d64

Observation 38d4dd40-061d-4311-b806-f3d669eca342 · outbound

This paper cites Cotdet: Affordance knowledge prompting for task driven object de- tection.

Training-free Generation of Temporally Consistent Rewards from VLMs Cotdet: Affordance knowledge prompting for task driven object de- tection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:47.390815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:41.990751Z digest=sha256:897fb4cb0c9db1e02740a8df3493286365d6b747baec66a66a2dc1267b21f443

Observation 97be814c-a937-4040-b3cd-c67d0f46dc28 · outbound

This paper cites AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter.

Training-free Generation of Temporally Consistent Rewards from VLMs AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.042018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.042018Z digest=sha256:031c56ac469008a9d9f68e226a6e08d2343f63c00995264efe48829516316fac

Observation 54bc3c62-7248-4cab-8587-cf1fc9edbe6f · outbound

This paper cites Real-World Offline Reinforcement Learning from Vision Language Model Feedback.

Training-free Generation of Temporally Consistent Rewards from VLMs Real-World Offline Reinforcement Learning from Vision Language Model Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.135720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.135720Z digest=sha256:628d7033cb8a19585dfc9fa1f233850581155b4690cac2a59f293bd79c9ef4da

Observation 27fd4f0c-f04d-427c-98df-383ab589ccce · outbound

This paper cites Code as reward: Empowering reinforcement learning with vlms.

Training-free Generation of Temporally Consistent Rewards from VLMs Code as reward: Empowering reinforcement learning with vlms

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:47.174138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:42.250922Z digest=sha256:41d880a82f88def2aa4ecf765289bfc5aa2734dfd84db968f67b693f39a8d7db

Observation bf4a0aaf-3f49-4f04-8f85-551259fb6f6f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Training-free Generation of Temporally Consistent Rewards from VLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.366889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.366889Z digest=sha256:83c0c292f629aad574b1dfb17a94888b265a80538493f5c4709984ea809e7fee

Observation 20d27e66-116d-4cbc-85aa-aea9cdc545fd · outbound

This paper cites Rl-vlm-f: Rein- forcement learning from vision language foundation model feedback.

Training-free Generation of Temporally Consistent Rewards from VLMs Rl-vlm-f: Rein- forcement learning from vision language foundation model feedback

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.922946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:42.528976Z digest=sha256:55d59aebde8fb7af0fb387f5ad3d9e675d0e8456f39557fa9a44098541d7bef1

Observation fb08866f-8e6e-4dcc-a15c-23a99cda2ebb · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

Training-free Generation of Temporally Consistent Rewards from VLMs Any-point Trajectory Modeling for Policy Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.614493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.614493Z digest=sha256:5dde9c09d580a91b411924492f3383ca9b1f1f386f56a0e8b8d2232a7d60de23

Observation 7dce3e65-a42f-40f7-8ed2-0869618216e2 · outbound

This paper cites RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation.

Training-free Generation of Temporally Consistent Rewards from VLMs RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.681720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.681720Z digest=sha256:c5fe4465d0e900dc0c3a81120805d4aa7b264a41aa7c3ec4e8f431fb12360549

Observation 63959c63-c49c-4b44-90e8-5d05e92514a1 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Training-free Generation of Temporally Consistent Rewards from VLMs Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:42.750529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:42.750529Z digest=sha256:1f002af5a0f36c063be0be7af4436516b6abdc807641cff61c0b1e5c5dce722a

Observation 7d982682-9648-4ece-baef-88452c82b263 · outbound

This paper cites Robot fine- tuning made easy: Pre-training rewards and policies for au- tonomous real-world reinforcement learning.

Training-free Generation of Temporally Consistent Rewards from VLMs Robot fine- tuning made easy: Pre-training rewards and policies for au- tonomous real-world reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.783281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:42.851403Z digest=sha256:4e1eb50608645e1cdb9c1cfe2f5f70efc5238da795e1811f5299ead9dac5afe2

Observation e2bb6c5d-f225-4e7b-806a-51239a3d6acf · outbound

This paper cites Par- ticle filters in latent space for robust deformable linear ob- ject tracking.

Training-free Generation of Temporally Consistent Rewards from VLMs Par- ticle filters in latent space for robust deformable linear ob- ject tracking

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.603324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:42.923587Z digest=sha256:1f85e7ff8e12156ea36f9f8a6667cfa114bd43d86b4c8fde7d5b199a2b5a5d05

Observation 4f28e9b1-6896-45fa-8085-92ec2fb70bd8 · outbound

This paper cites Sornet: Spatial object-centric representations for se- quential manipulation.

Training-free Generation of Temporally Consistent Rewards from VLMs Sornet: Spatial object-centric representations for se- quential manipulation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.378775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:42.987023Z digest=sha256:67c29da9b0a6f1adca045c2e0e377a0d3cbefda11a8eaeb2f2013818a5638937

Observation 218006e5-280a-4e00-837c-38e8e384aa60 · outbound

This paper cites Robopoint: A vision-language model for spatial affordance prediction in robotics.

Training-free Generation of Temporally Consistent Rewards from VLMs Robopoint: A vision-language model for spatial affordance prediction in robotics

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.158520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:43.078554Z digest=sha256:f5b291b96169a356c83e72c818d442803ef0a3a58820ff786847f9e7798563b2

Observation c571f2e5-095b-4f16-bfaf-fea0bb00de83 · outbound

This paper cites Sam-e: Leveraging visual foundation model with sequence imitation for embodied ma- nipulation.

Training-free Generation of Temporally Consistent Rewards from VLMs Sam-e: Leveraging visual foundation model with sequence imitation for embodied ma- nipulation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:46.014193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:43.151441Z digest=sha256:32f24a5aeaeb3f8c0455305900138d26b1e900605fe6963312cd2a08601afd2c

Observation 429af5d3-1448-4a2c-a090-499261c9dcb6 · outbound

This paper cites MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation.

Training-free Generation of Temporally Consistent Rewards from VLMs MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:43.218177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:43.218177Z digest=sha256:8ea7ff271d9c97d134d7bf3c9a1ffd5adac9286afa2e0a7925dad2610d588e3e

Observation e2bde3eb-7b89-43d8-ac34-52fb8082f21b · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

Training-free Generation of Temporally Consistent Rewards from VLMs TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:43.290054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:43.290054Z digest=sha256:357495ac5b0a19e2605e004ee13a5d1be973530c392ff9951f3a8841e66c5dd3

Observation e9e5d94c-0ddc-42ff-9e4a-a575fa5fcfce · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:45.846680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:43.353180Z digest=sha256:d2e4ef4376dcf3266c3a57497cb7c65a11fc0e79949fd2220b5c3a28bbfaafcd

Observation 6a8e0e12-be39-4bf6-86ce-c1690d17dbb5 · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:45.286813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:43.546848Z digest=sha256:b9564d3a0cb9dcbd2bc75e09833d581fd869649d71a30895d2e5076798ea8d5e

Observation 95211ab0-e2b3-4a49-8da8-7f736371a782 · outbound

This paper cites Figure 11.

Training-free Generation of Temporally Consistent Rewards from VLMs Figure 11

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:45.081039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:43.639443Z digest=sha256:49e2462b375682b554c8562171047cc6477cc0e65ee5fa7565b7d3c6bc9f29a3

Observation a9656fe6-cc2f-4b2f-bdfb-2fb0cc0ead1c · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:44.831500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:43.723254Z digest=sha256:ae68329c4c3cf91da1e2caf1d895f31292b059a9ae44090c4d0e457a53440e4f

Observation d2b99751-8743-4fa9-a03e-4a64aaf32e54 · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:45.446685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:43.790046Z digest=sha256:4c669c1120d260cdf02a9f59bb6a9104647c6d817954c28aa3e7ee15f87280d2

Observation db95df95-ace1-4a26-899d-51c16c8dae9c · outbound

This paper cites an unresolved cited work.

Training-free Generation of Temporally Consistent Rewards from VLMs Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:45.620439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:43.909897Z digest=sha256:2cb8cd8e3f41e62f33e0ab53a323173f02ddb8304350d76e392bcb10e7e83e7e

Observation 4fe6474c-0e82-4c38-a003-dcb777e0049d · outbound

This paper cites Completion Status Identification of Sub-goals: System prompt: Detailed in Fig.

Training-free Generation of Temporally Consistent Rewards from VLMs Completion Status Identification of Sub-goals: System prompt: Detailed in Fig

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:44.614948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:47:44.014826Z digest=sha256:02f5f70a47008325ff429de9bfe503b47b717edcb7fe63f23f944b3f7b7b6bd7

Pith citing papers

No inbound Pith citation observations are available.