Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T10:53:06.346339Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2507.23391.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T10:53:06.346339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d86a3918-632c-41cb-96db-7e84a500fd33 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Grandmaster level in starcraft ii using multi-agent reinforcement learning,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 14bd5ec2-334b-4a45-a8ee-d429ab54d3de · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Mastering the game of stratego with model-free multiagent reinforcement learning,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 17d0a040-5158-448d-b2c2-46f60e144496 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Autonomous navigation of stratospheric balloons using reinforcement learning,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 71e3c09d-2e49-45c0-b2e5-7d3349fd3205 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Champion-level drone racing using deep reinforce- ment learning,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4568d228-adea-4d08-b181-4542ab7bbd6e · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Scalable deep reinforcement learning for vision-based robotic manipulation,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 37c59ca1-485b-49f0-bf85-3311484cbdeb · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Hindsight goal ranking on replay buffer for sparse reward environment,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 168a47ab-deec-41d2-b009-e56515b5f845 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Towards human-level bimanual dexterous manipulation with reinforcement learning,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 50c51c11-9492-47af-9a70-88c7bf7445b3 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Inverse reward design,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 72ac7b64-fd7e-47fa-a629-f32c811466c5 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Defining and characterizing reward gaming,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 33304b4d-4ea2-4024-a569-c351f393f74e · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33e373fa-455c-41a4-b86f-ae37427f8b55 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Hear: Hearing enhanced audio response for video-grounded dialogue,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9861685b-8ecc-4ebf-84cd-ae559635bbee · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ff0ba41-1dde-4a6d-aaf5-81e551420daa · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Openvla: An open- source vision-language-action model,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 90785889-ab61-4ee7-baca-7740c7e856ce · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Code as reward: Empowering reinforcement learning with vlms,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 85304670-8ea2-448d-9050-bcfb8f5d74db · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Guiding pretraining in reinforcement learning with large language models,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 714b9b62-b76f-476f-b886-341a23db2840 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Bootstrap your own skills: Learning to solve new tasks with large language model guidance,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5e4493de-196c-412e-8578-dcd5b8cb6686 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Code as policies: Language model programs for embodied control,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c85d74e0-5ce3-40af-a9b5-ffa6c0a81b6f · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Progprompt: Generating situ- ated robot task plans using large language models,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5d23bd93-0fc2-44dc-9464-340108aa7af3 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Learning transferable visual models from natural language supervision,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 70a94aae-ba42-47f9-9077-52c2af4d89e6 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Roboclip: One demonstration is enough to learn robot policies,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8eef3b20-dae8-4971-b742-c7fc13caa125 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Vision-language models are zero-shot reward models for reinforce- ment learning,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fbf8ed2d-8284-418f-99fa-3dc39fbebcaf · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Rl-vlm-f: Reinforcement learning from vision language foundation model feedback,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 550f14f4-b8f1-4d14-b23e-eb9092bb1a0e · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Language to rewards for robotic skill synthesis,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 877111fb-e0e4-422f-a5cf-5b3f9e168413 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Text2reward: Automated dense reward function generation for reinforcement learning,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0733e202-e99f-439a-a7b8-d8c258473669 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Real- world offline reinforcement learning from vision language model feedback,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 43179de3-2947-4c83-a0ea-d96d4e7f53b8 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Deep reinforcement learning from human preferences,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation be018ad3-eb4a-4589-9da2-1aaee9add368 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling B-pref: Benchmark- ing preference-based reinforcement learning,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ef05fdd8-34a2-4753-8d8e-5f56fad7ed87 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Inverse preference learning: Preference- based rl without a reward function,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2e30d69f-aba6-4241-a89e-b4e115485c2e · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Information- theoretic text hallucination reduction for video-grounded dialogue,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f4628f95-d6db-40ca-9056-8bea09610fe8 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79335cdf-3635-4a51-9b77-e07e75cdfb78 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling ” task success
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8c522c47-4cd2-4e5e-b092-a9c183aac69f · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Non-markovian reward modelling from trajectory labels via interpretable multiple instance learning,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 72d88e7a-fda6-49a0-a016-d2e5dccc8803 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Preference transformer: Modeling human preferences using transformers for rl,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8dbda9e9-d510-449b-b719-610c18c56087 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18443a1e-ce39-4706-8a92-b57986a1c008 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Contrastive preference learning: learning from human feedback without rl,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e0db7229-ff58-482d-802c-b46eae455e41 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bbe84bfb-b9ae-474f-a2bc-5ed0905421cc · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Minedojo: Building open- ended embodied agents with internet-scale knowledge,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 42481d8c-6090-473e-8dd8-f976c4c9b6ac · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Liv: Language-image representations and rewards for robotic control,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2795730f-c34c-427c-9793-7c588fb38b78 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Enhancing rating-based reinforcement learning to effectively leverage feedback from large vision-language models,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e63164c5-bc88-42d4-a54d-b93b68b8e753 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Alvinn: An autonomous land vehicle in a neural network,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d7dab82b-2566-4c20-bbe7-324be4c90642 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling RT-1: Robotics Transformer for Real-World Control at Scale
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8707109c-e502-4633-a6ed-39610003a681 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Interactive language: Talking to robots in real time,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 01b44d5b-b02c-4643-b407-02bed96ab2ac · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Predictive coding for decision transformer,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b0511e0f-0019-4d67-a2d8-c5378948887e · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Algorithms for inverse reinforcement learning.,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 05f636e7-57df-4bfd-804d-f60467f9aa22 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Generative adversarial imitation learning,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5c874bb1-a9d9-4080-a848-c047f95b7357 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Nonlinear inverse reinforcement learning with gaussian processes,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fc0123b2-7a7a-4255-99e3-5004f8857d7b · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Guided cost learning: Deep inverse optimal control via policy optimization,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation de2a873b-0c9f-4990-8f7a-3547bfe089b1 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Learning robust rewards with adversarial inverse reinforcement learning,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation db2a8b19-33a6-4d4a-b545-3af8aeb66942 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Confidence-aware imitation learning from demonstrations with varying optimality,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3cf7a57e-9f41-4a69-b1b7-6b6cdc7fdc91 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5b55927f-a36d-4880-84e3-79efe33d8432 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Denoising diffusion probabilistic models,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 39241caa-1c95-4fff-a31c-5ed6ea7859dc · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Mdsgen: Fast and efficient masked diffusion temporal-aware transformers for open-domain sound generation,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9f947a2b-236d-4fd3-9404-54abcfba526a · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Taro: Timestep-adaptive repre- sentation alignment with onset-aware conditioning for synchronized video-to-audio synthesis,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5e92706e-7a7b-4323-a8cf-85ed86701f44 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Diffusion reward: Learning rewards via conditional video diffusion,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cb140fe9-36e6-4666-9332-607cb67ebe4f · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Di- rect preference-based policy optimization without reward modeling,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 68b9331c-313b-4c30-9f7f-21f90f6926a2 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Atari-head: Atari human eye- tracking and demonstration dataset,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 87eedce5-eaf1-4e6b-9fe1-fceda5a9af27 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Scanet: Scene complexity aware network for weakly-supervised video moment retrieval,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7e55aa6e-935b-4cbe-902f-f1333c2e366f · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Learning deep networks from noisy labels with dropout regularization,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 10c01c1f-f6a6-456c-ad6e-3b25919a5cb2 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Compressing features for learning with noisy labels,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b11d135b-c4d3-4f63-bb9c-8fc8edefdc85 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling R3m: A universal visual representation for robot manipulation,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 21a841f6-0ff4-45d1-a87a-9c64886d5b75 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Offline reinforcement learning with implicit q-learning,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 071a7a7b-21da-46ea-a460-94122014cc96 · outbound
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling Rime: Robust preference-based reinforcement learning with noisy preferences,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.