Pith. sign in

Paper Citation Record · LEDGER

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling

As of 10 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2507.23391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.23391 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:53:06.346339Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy57
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d86a3918-632c-41cb-96db-7e84a500fd33 · outbound

This paper cites Grandmaster level in starcraft ii using multi-agent reinforcement learning,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Grandmaster level in starcraft ii using multi-agent reinforcement learning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.321112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.088233Z digest=sha256:7b17aa8f04a94aac2831d6a7225fa888afec1e6f49bcbffb552fa41487e089d2

Observation 14bd5ec2-334b-4a45-a8ee-d429ab54d3de · outbound

This paper cites Mastering the game of stratego with model-free multiagent reinforcement learning,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Mastering the game of stratego with model-free multiagent reinforcement learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.303375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.093365Z digest=sha256:1318284b003ccf8df047bef74f783291fb68ab2bc5d333b9588033d992dc2462

Observation 17d0a040-5158-448d-b2c2-46f60e144496 · outbound

This paper cites Autonomous navigation of stratospheric balloons using reinforcement learning,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Autonomous navigation of stratospheric balloons using reinforcement learning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.286207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.098840Z digest=sha256:a327350149609cc9d1085811e922401912823f0b3e4a9280bd12d277ac6133d0

Observation 71e3c09d-2e49-45c0-b2e5-7d3349fd3205 · outbound

This paper cites Champion-level drone racing using deep reinforce- ment learning,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Champion-level drone racing using deep reinforce- ment learning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.271442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.103807Z digest=sha256:f0a0f2efd82d2ebcfc1dfc8182c1d8c578c5c960cf6528f43a8d6a33c7d42a2c

Observation 4568d228-adea-4d08-b181-4542ab7bbd6e · outbound

This paper cites Scalable deep reinforcement learning for vision-based robotic manipulation,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Scalable deep reinforcement learning for vision-based robotic manipulation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.253460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.109768Z digest=sha256:ca909d727b8045c7bf6ee6873b2ca75e1f6eece2801185322710b56370fdb08f

Observation 37c59ca1-485b-49f0-bf85-3311484cbdeb · outbound

This paper cites Hindsight goal ranking on replay buffer for sparse reward environment,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Hindsight goal ranking on replay buffer for sparse reward environment,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.226778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.114453Z digest=sha256:b74249d2de535efb6c8d0cace73cd52c10b6395828b2db3153e12606ca750d46

Observation 168a47ab-deec-41d2-b009-e56515b5f845 · outbound

This paper cites Towards human-level bimanual dexterous manipulation with reinforcement learning,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Towards human-level bimanual dexterous manipulation with reinforcement learning,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.204451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.119415Z digest=sha256:73b271029cf0eb5db5acffb785feb4876bf2d99215f727f8abe25845a67118ba

Observation 50c51c11-9492-47af-9a70-88c7bf7445b3 · outbound

This paper cites Inverse reward design,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Inverse reward design,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.189636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.123667Z digest=sha256:e6df686e3955edbf36356e033f646c03d829480f7bd4208fccbb2be902071a11

Observation 72ac7b64-fd7e-47fa-a629-f32c811466c5 · outbound

This paper cites Defining and characterizing reward gaming,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Defining and characterizing reward gaming,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.175458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.127814Z digest=sha256:f5587b5d9bac7271f0d546517f7eaab8e2d17147a94ccfe48acbe235083c54a5

Observation 33304b4d-4ea2-4024-a569-c351f393f74e · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T10:53:06.131627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:53:06.131627Z digest=sha256:1dd361bad0b74fd0d527359dff5f0164d212f0e41c513cbcf46950ef56fc0513

Observation 33e373fa-455c-41a4-b86f-ae37427f8b55 · outbound

This paper cites Hear: Hearing enhanced audio response for video-grounded dialogue,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Hear: Hearing enhanced audio response for video-grounded dialogue,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.161546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.136559Z digest=sha256:30361eb2c52dbc5f283907bfaea90437205c91423d827c2c7613f8439cbe9b07

Observation 9861685b-8ecc-4ebf-84cd-ae559635bbee · outbound

This paper cites AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T10:53:06.140763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:53:06.140763Z digest=sha256:b4406e59d8e7802ff29220aa5185fe2b1ebad56e830487e8e1a95a74ee2d791b

Observation 3ff0ba41-1dde-4a6d-aaf5-81e551420daa · outbound

This paper cites Openvla: An open- source vision-language-action model,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Openvla: An open- source vision-language-action model,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.148207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.145585Z digest=sha256:76b8797392f74953f3c96ca1475db6508eadd9473dc97fc697ae1ec7d5feacbf

Observation 90785889-ab61-4ee7-baca-7740c7e856ce · outbound

This paper cites Code as reward: Empowering reinforcement learning with vlms,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Code as reward: Empowering reinforcement learning with vlms,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.133944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.150078Z digest=sha256:a83f1c44474abeabbc0cd2f4223f506b664ce78405198266711754d1ed77bb49

Observation 85304670-8ea2-448d-9050-bcfb8f5d74db · outbound

This paper cites Guiding pretraining in reinforcement learning with large language models,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Guiding pretraining in reinforcement learning with large language models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.119511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.154046Z digest=sha256:42a1f793976e2d35aa3ca661bdf39f36fae0cc0d4b15eb038586890effedf2b6

Observation 714b9b62-b76f-476f-b886-341a23db2840 · outbound

This paper cites Bootstrap your own skills: Learning to solve new tasks with large language model guidance,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Bootstrap your own skills: Learning to solve new tasks with large language model guidance,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.105934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.158153Z digest=sha256:f3a14628f5de4005a5c760ea58009666e5b6af76c78e60cb0f0449585c2a1c74

Observation 5e4493de-196c-412e-8578-dcd5b8cb6686 · outbound

This paper cites Code as policies: Language model programs for embodied control,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Code as policies: Language model programs for embodied control,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.092568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.162053Z digest=sha256:6ce39face52525bd37efe273b44703e9f0cc8cfb688eb7a69feddbcf6f116b61

Observation c85d74e0-5ce3-40af-a9b5-ffa6c0a81b6f · outbound

This paper cites Progprompt: Generating situ- ated robot task plans using large language models,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Progprompt: Generating situ- ated robot task plans using large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.079466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.166100Z digest=sha256:46e68fb8fded08b566b09cc5f68083eb4b2e10f44c783146ec17409499d9f70d

Observation 5d23bd93-0fc2-44dc-9464-340108aa7af3 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Learning transferable visual models from natural language supervision,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.066556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.170123Z digest=sha256:c45e082bd7cea912ffedc5cec915ead47999ab7fc2d71f05113cf250193b659c

Observation 70a94aae-ba42-47f9-9077-52c2af4d89e6 · outbound

This paper cites Roboclip: One demonstration is enough to learn robot policies,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Roboclip: One demonstration is enough to learn robot policies,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.052754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.173998Z digest=sha256:7379b81c3cdc0b84d12ebc39df7b4f626d16b4ebd08ec7f0370ac65a3f263509

Observation 8eef3b20-dae8-4971-b742-c7fc13caa125 · outbound

This paper cites Vision-language models are zero-shot reward models for reinforce- ment learning,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Vision-language models are zero-shot reward models for reinforce- ment learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.039310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.178420Z digest=sha256:c842a642edd896daea5b8f4673a754f3988cae66260e4e0285fc736816a35dc6

Observation fbf8ed2d-8284-418f-99fa-3dc39fbebcaf · outbound

This paper cites Rl-vlm-f: Reinforcement learning from vision language foundation model feedback,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Rl-vlm-f: Reinforcement learning from vision language foundation model feedback,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.025349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.182223Z digest=sha256:72bb5ab0b821eba3daf5f30d3b591ff88dbd7c663b1e9fe53908d14baf6d8d3f

Observation 550f14f4-b8f1-4d14-b23e-eb9092bb1a0e · outbound

This paper cites Language to rewards for robotic skill synthesis,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Language to rewards for robotic skill synthesis,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:07.007379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.186024Z digest=sha256:f9c78afced9b0727656438605844e002cdbc985752f3399c161f79942788a173

Observation 877111fb-e0e4-422f-a5cf-5b3f9e168413 · outbound

This paper cites Text2reward: Automated dense reward function generation for reinforcement learning,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Text2reward: Automated dense reward function generation for reinforcement learning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.992559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.190298Z digest=sha256:af95da575438722f8ab399c3b11dac81cb7d814113b9d0b5902a4d21a170b396

Observation 0733e202-e99f-439a-a7b8-d8c258473669 · outbound

This paper cites Real- world offline reinforcement learning from vision language model feedback,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Real- world offline reinforcement learning from vision language model feedback,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.976824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.194150Z digest=sha256:9623d7c0d3258080d3d5c4be6274f1916e1e9dbb30dcc8dee6cef8711898ee84

Observation 43179de3-2947-4c83-a0ea-d96d4e7f53b8 · outbound

This paper cites Deep reinforcement learning from human preferences,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Deep reinforcement learning from human preferences,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.961152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.198326Z digest=sha256:6f4727fbaceeb91dd1f10b5d957a7f9be56997868f1863754c363721e7c48b00

Observation be018ad3-eb4a-4589-9da2-1aaee9add368 · outbound

This paper cites B-pref: Benchmark- ing preference-based reinforcement learning,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling B-pref: Benchmark- ing preference-based reinforcement learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.946274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.202410Z digest=sha256:6fe631d1350176d44bc1077afd7972dbf959bef4ecf978ac40e6619f24003114

Observation ef05fdd8-34a2-4753-8d8e-5f56fad7ed87 · outbound

This paper cites Inverse preference learning: Preference- based rl without a reward function,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Inverse preference learning: Preference- based rl without a reward function,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.930760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.206140Z digest=sha256:0958a8ac6cd4e9bca32f572c57f40d3c8e6f8f718ab08712c2e470d36ed9feef

Observation 2e30d69f-aba6-4241-a89e-b4e115485c2e · outbound

This paper cites Information- theoretic text hallucination reduction for video-grounded dialogue,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Information- theoretic text hallucination reduction for video-grounded dialogue,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.915791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.210306Z digest=sha256:676e0e836ae698a3fd15c72e99d713d50f4faad2dbc80f65864057fa57ba5da4

Observation f4628f95-d6db-40ca-9056-8bea09610fe8 · outbound

This paper cites Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T10:53:06.214265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:53:06.214265Z digest=sha256:71db1ea7a31329faa2a46379de99081367e5e2a3c730b53c2c7b9d12070af8cd

Observation 79335cdf-3635-4a51-9b77-e07e75cdfb78 · outbound

This paper cites ” task success.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling ” task success

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.902589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.218449Z digest=sha256:461606ea31a30a85255a22a9d9957bb5b72980851a363468907285fb0f836e9b

Observation 8c522c47-4cd2-4e5e-b092-a9c183aac69f · outbound

This paper cites Non-markovian reward modelling from trajectory labels via interpretable multiple instance learning,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Non-markovian reward modelling from trajectory labels via interpretable multiple instance learning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.887933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.222272Z digest=sha256:43d30a7c916e45aaf7b17a4fdb96b577bbc5952d18aa4a32d99a438ec46a8d3e

Observation 72d88e7a-fda6-49a0-a016-d2e5dccc8803 · outbound

This paper cites Preference transformer: Modeling human preferences using transformers for rl,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Preference transformer: Modeling human preferences using transformers for rl,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.874163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.226091Z digest=sha256:b3d3ac3558dbf76ac4c6313ffdef5b53f8bfa63b6b0e46acb73ff092a78467bb

Observation 8dbda9e9-d510-449b-b719-610c18c56087 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T10:53:06.230845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:53:06.230845Z digest=sha256:befe628f4c6254643dfca0fc4887b827ad4e3c2374e16138f16e4585a8220e26

Observation 18443a1e-ce39-4706-8a92-b57986a1c008 · outbound

This paper cites Contrastive preference learning: learning from human feedback without rl,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Contrastive preference learning: learning from human feedback without rl,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.859858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.235540Z digest=sha256:49f16acfaf38c0b0d89d97a56757a826be7923d2be32a76aade954d637b778e0

Observation e0db7229-ff58-482d-802c-b46eae455e41 · outbound

This paper cites Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.846472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.240448Z digest=sha256:a9f9dfa5400002cfc90c720499bb08b48c1dbda8b79d624401dd6cb8d11c8571

Observation bbe84bfb-b9ae-474f-a2bc-5ed0905421cc · outbound

This paper cites Minedojo: Building open- ended embodied agents with internet-scale knowledge,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Minedojo: Building open- ended embodied agents with internet-scale knowledge,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.831007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.244892Z digest=sha256:ace00852334bca91ced08a068e89940ae9798b7532849471faa4d77d41bd445a

Observation 42481d8c-6090-473e-8dd8-f976c4c9b6ac · outbound

This paper cites Liv: Language-image representations and rewards for robotic control,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Liv: Language-image representations and rewards for robotic control,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.815934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.248963Z digest=sha256:b51865cb9b2ce1566bd3691a7d4e2c6da44b5fabd0a0a84bec38211878183aab

Observation 2795730f-c34c-427c-9793-7c588fb38b78 · outbound

This paper cites Enhancing rating-based reinforcement learning to effectively leverage feedback from large vision-language models,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Enhancing rating-based reinforcement learning to effectively leverage feedback from large vision-language models,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.801308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.253129Z digest=sha256:d5bb105439e8011af165436100361a8d667409d60aa786391f771734142a3cc6

Observation e63164c5-bc88-42d4-a54d-b93b68b8e753 · outbound

This paper cites Alvinn: An autonomous land vehicle in a neural network,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Alvinn: An autonomous land vehicle in a neural network,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.785403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.256898Z digest=sha256:d55fd806e9578bfa1ffc0233e4fc4d3d63a1e422b612af14190403a1f40e74c4

Observation d7dab82b-2566-4c20-bbe7-324be4c90642 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling RT-1: Robotics Transformer for Real-World Control at Scale

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T10:53:06.260783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:53:06.260783Z digest=sha256:4c80304de113bbc535be24975ca9d5a6180338fff7b56f38c0478201b0221a15

Observation 8707109c-e502-4633-a6ed-39610003a681 · outbound

This paper cites Interactive language: Talking to robots in real time,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Interactive language: Talking to robots in real time,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.771071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.265314Z digest=sha256:122a096f50f68d2768e4e81cdb4e8750debdf925fb24c17ff9154b14b674d767

Observation 01b44d5b-b02c-4643-b407-02bed96ab2ac · outbound

This paper cites Predictive coding for decision transformer,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Predictive coding for decision transformer,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.757046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.269155Z digest=sha256:cd93b6b08a2893f900b6dcce14dd288dc590ec32ae3f636e30bc385debe07817

Observation b0511e0f-0019-4d67-a2d8-c5378948887e · outbound

This paper cites Algorithms for inverse reinforcement learning.,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Algorithms for inverse reinforcement learning.,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.743708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.273130Z digest=sha256:16f206b19c1f8ac45e11d83025149f5e957f24e44ad86024b2e926f12f8e30b1

Observation 05f636e7-57df-4bfd-804d-f60467f9aa22 · outbound

This paper cites Generative adversarial imitation learning,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Generative adversarial imitation learning,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.729304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.277216Z digest=sha256:f00c3219d1866b814b6f3f2378d2ddaecfe83f0423d47baf49e243fcff652b1c

Observation 5c874bb1-a9d9-4080-a848-c047f95b7357 · outbound

This paper cites Nonlinear inverse reinforcement learning with gaussian processes,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Nonlinear inverse reinforcement learning with gaussian processes,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.712319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.280942Z digest=sha256:42158fe3d9d9c9fd2b1ebcc636bb12c4190141a9364ee3d32a30654a3d36834d

Observation fc0123b2-7a7a-4255-99e3-5004f8857d7b · outbound

This paper cites Guided cost learning: Deep inverse optimal control via policy optimization,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Guided cost learning: Deep inverse optimal control via policy optimization,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.697777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.285120Z digest=sha256:8e21050faec7b56f4cd07fdb3d6ab195f35f07ca3d5af704fd6e4bba5f8c8658

Observation de2a873b-0c9f-4990-8f7a-3547bfe089b1 · outbound

This paper cites Learning robust rewards with adversarial inverse reinforcement learning,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Learning robust rewards with adversarial inverse reinforcement learning,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.681468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.289137Z digest=sha256:493718365a2f441ff5cdc2006e9cf370b65be8ab119cc0a58b24591566ca25c1

Observation db2a8b19-33a6-4d4a-b545-3af8aeb66942 · outbound

This paper cites Confidence-aware imitation learning from demonstrations with varying optimality,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Confidence-aware imitation learning from demonstrations with varying optimality,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.656579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.293287Z digest=sha256:11f82a640cf2662ea8ee9bcd9b7469227798db599c0edfd4583f9f44a1ad195d

Observation 3cf7a57e-9f41-4a69-b1b7-6b6cdc7fdc91 · outbound

This paper cites Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.641140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.297923Z digest=sha256:b6ee32be3d3990e03cb9724dec481ef837e6133448259205e1bb39bbc670f268

Observation 5b55927f-a36d-4880-84e3-79efe33d8432 · outbound

This paper cites Denoising diffusion probabilistic models,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Denoising diffusion probabilistic models,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.626234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.301997Z digest=sha256:3fc89f91b6961db74eff6c2cbe630a7aee9c18a9e1cfddb1e679719c256b7258

Observation 39241caa-1c95-4fff-a31c-5ed6ea7859dc · outbound

This paper cites Mdsgen: Fast and efficient masked diffusion temporal-aware transformers for open-domain sound generation,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Mdsgen: Fast and efficient masked diffusion temporal-aware transformers for open-domain sound generation,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.611140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.306319Z digest=sha256:4f0c75994f227e15f5c17ee72d38aa6398f95c7d1c85778a62bb4a7e36410014

Observation 9f947a2b-236d-4fd3-9404-54abcfba526a · outbound

This paper cites Taro: Timestep-adaptive repre- sentation alignment with onset-aware conditioning for synchronized video-to-audio synthesis,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Taro: Timestep-adaptive repre- sentation alignment with onset-aware conditioning for synchronized video-to-audio synthesis,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.593396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.310276Z digest=sha256:9a7916653c8ed8e48b95e43840ed619d37e7f25307c42ce4f5cd09fc5e8776e0

Observation 5e92706e-7a7b-4323-a8cf-85ed86701f44 · outbound

This paper cites Diffusion reward: Learning rewards via conditional video diffusion,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Diffusion reward: Learning rewards via conditional video diffusion,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.578865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.314133Z digest=sha256:0280579e48c030b8aea4a36a2c99959119bcc8b9b4b9244d0448c8db1d8d9205

Observation cb140fe9-36e6-4666-9332-607cb67ebe4f · outbound

This paper cites Di- rect preference-based policy optimization without reward modeling,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Di- rect preference-based policy optimization without reward modeling,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.563052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.318420Z digest=sha256:a4584ebd8320e4ad54225182dbe9513febfe47f04ef0c578339e9b8922873bf6

Observation 68b9331c-313b-4c30-9f7f-21f90f6926a2 · outbound

This paper cites Atari-head: Atari human eye- tracking and demonstration dataset,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Atari-head: Atari human eye- tracking and demonstration dataset,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.549341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.322397Z digest=sha256:5f48cc36e3c2d89edd27261752181f64e577fc578d57c5555fec3c7ecc443ee8

Observation 87eedce5-eaf1-4e6b-9fe1-fceda5a9af27 · outbound

This paper cites Scanet: Scene complexity aware network for weakly-supervised video moment retrieval,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Scanet: Scene complexity aware network for weakly-supervised video moment retrieval,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.535055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.326259Z digest=sha256:99557708e0021273c0eec12505e875ccd395ff440aabf16a27aefd744cc7e493

Observation 7e55aa6e-935b-4cbe-902f-f1333c2e366f · outbound

This paper cites Learning deep networks from noisy labels with dropout regularization,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Learning deep networks from noisy labels with dropout regularization,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.519618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.330269Z digest=sha256:8323e8d6029b9009487d8e52718365f6933bb94fdd17c3550214138b8e756161

Observation 10c01c1f-f6a6-456c-ad6e-3b25919a5cb2 · outbound

This paper cites Compressing features for learning with noisy labels,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Compressing features for learning with noisy labels,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.505352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.334207Z digest=sha256:d757f22db019aac5cf4a88ff3c9c090288198769a46bdb927e7c1d9d89565f44

Observation b11d135b-c4d3-4f63-bb9c-8fc8edefdc85 · outbound

This paper cites R3m: A universal visual representation for robot manipulation,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling R3m: A universal visual representation for robot manipulation,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.489708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.338084Z digest=sha256:1bf3cb63402b462f150ca33573226b2dd60496a25a96d8838e65e78f5e0133c6

Observation 21a841f6-0ff4-45d1-a87a-9c64886d5b75 · outbound

This paper cites Offline reinforcement learning with implicit q-learning,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Offline reinforcement learning with implicit q-learning,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.474911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.342348Z digest=sha256:cf39536a8bfc79be58b8ef21acefa214d78bffdc7d5392aa3d202ec20f857f7b

Observation 071a7a7b-21da-46ea-a460-94122014cc96 · outbound

This paper cites Rime: Robust preference-based reinforcement learning with noisy preferences,.

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Rime: Robust preference-based reinforcement learning with noisy preferences,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:53:06.460060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:53:06.346339Z digest=sha256:57e6d31d093fe78d84fdc8bcb44871e9feda49aeff782d415265b835b608fbc0

Pith citing papers

No inbound Pith citation observations are available.