Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:11:33.916383Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2509.06759.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:11:33.916383Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T10:06:15.623188Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T03:26:28.843655Z
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0b7ddbe5-6805-4d44-8d87-936c0c175228 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Visual instruction tuning,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82d881f1-e063-4db4-87d4-7c06c0718ac1 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization MiniGPT-4: Enhancing vision-language understanding with advanced large language models,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a9705884-29bd-481d-ad90-0883b4698e2a · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization FuRL: visual-language models as fuzzy rewards for reinforcement learning,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 285986a1-08cd-473a-b33e-adbfa20b63a4 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Direct preference optimization: your language model is secretly a reward model,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 937d651a-f762-4345-9648-4634aaeccb1f · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Proximal Policy Optimization Algorithms
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81e810d6-b6a1-4164-b624-1301cfb1878b · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Deep reinforcement learning for cyber security,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2530fe82-4bf8-404e-b646-29e7b055609f · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eb4008c-f784-4649-a0cc-4dd41817ff9d · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96de6b97-b3f9-44e5-8037-ad88ddfeb846 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization GPT-4V(ision) System Card,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f32b6995-f3cb-4e21-a4d0-c354ea1c84f7 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b2055ae-7013-4777-aa96-0396d55ff61c · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization The Claude 3 model family: Opus, Sonnet, Haiku,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4d1a607-5086-4fe1-8c10-dd42cd967546 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7afd6ff-f37d-48fc-8971-c23572d2a6f2 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06c642ba-e757-4f1a-b582-dd08acd688e4 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Qwen2.5-VL Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dea6216-c277-41e8-8d9c-deff99a9c82c · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05d60a77-9a45-47f1-aa8b-e5d7ef872968 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Aligning large multimodal models with factually augmented RLHF,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5515cda2-59b6-4b31-9a73-4a4ed8eef351 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Fusing pre-trained language models with multimodal prompts through reinforcement learning,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eca955ef-ec96-4f6a-bc04-9842a5ae3694 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Learning transferable visual models from natural language supervision,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2725aca8-5c7c-48f3-8fbf-57cd903c3e8f · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Improving Vision-Language-Action Model with Online Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 584d7002-35ed-412f-885a-9e1b337db0e4 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Fine-tuning large vision-language models as decision-making agents via reinforcement learning,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9dd998c8-ec99-4eb2-951b-5774943f28c6 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization VLFeedback: A large-scale AI feedback dataset for large vision-language models alignment,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e71b9a7-b5a1-432e-9845-65efc072eef0 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d62c896b-9f3a-4713-8d08-1799de8b65e5 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73f1b11e-3aba-4ab6-9a34-d35332c2bece · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Training language models to follow instructions with human feedback,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd0286eb-b055-4f36-99dd-9cb8bb277b00 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization RLHF-V: Towards trustworthy MLLMs via behavior alignment from fine-grained correctional human feedback,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ed9281b-9ea0-449d-809a-5428edc78ccd · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization RLAIF-V: Aligning MLLMs through open- source AI feedback for super GPT-4V trustworthiness,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c0e0dba-87fe-4724-8c9b-7450768ae793 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d9d970f-7b5a-4ed0-ba28-4912501b0c8b · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Aligning modalities in vision large language models via preference fine-tuning,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 98be5315-ceb8-4cbd-9ca8-eb2959294c31 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Re-Align: Aligning vision language models via retrieval-augmented direct preference optimization,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44d826cd-ca6d-4d51-a385-502c419dcf59 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Enhancing visual-language modality alignment in large vision language models via self-improvement,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 340830de-37de-418f-a89a-309ae2e27376 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization VILA: On pre-training for visual language models,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08d7f4a5-692d-40ce-a274-cc75b88179e0 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb03875e-d4b2-4db5-be8d-34f9a673557b · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization DRESS: Instructing large vision-language models to align and interact with humans via natural language feedback,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4f1bbaf7-89cd-4365-bcb9-edf78b7b875c · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Active learning for vision-language models,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 947346f9-f031-49f6-8893-3e196311fc20 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization S-CLIP: Semi-supervised vision- language learning using few specialist captions,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 39cd1753-ca90-4451-b2a6-3e7b88b37ceb · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization RL-VLM-F: reinforcement learning from vision language foundation model feedback,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f86d631-de35-4c62-b3b9-ab9639bc0c06 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50ab2a2b-223e-4db1-a473-b074a11bb0f2 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Text-to-decision agent: Offline meta-reinforcement learning from natural language supervision,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 275b609a-a0c7-4d8e-87f6-14be0bb1e33e · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Zero-shot model-based reinforcement learning using large language models,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81da2cfc-2953-4d37-9707-0803e844cfca · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb608950-511f-4461-ab83-d883d128b16e · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Open-ended VQA benchmarking of vision-language models by exploiting classification datasets and their semantic hierarchy,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b37a8033-2b5a-4f9b-b754-91d30b904049 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization VLP: Vision-Language Preference Learning for Embodied Manipulation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c0650d6e-0169-4640-a19c-fd7cc2086343 · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Multi-agent deep reinforcement learning with human strategies,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7336f9a6-1f18-42b4-a64b-401378996dcb · outbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization ETA: Evaluating then align- ing safety of vision language models at inference time,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4c65fd8-781c-4f7a-88f2-3b0a19ee93ef · inbound
Beyond the Literal: Decomposing Pragmatic Intent in Multimodal Meme Understanding Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.