Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:51:32.247216Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 5 inbound Pith citation observations for arXiv:2505.20777.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:51:32.247216Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T18:40:14.013017Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T22:06:16.638707Z
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f73ea394-c64f-4ec3-960f-677426875167 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Vision-language models for vision tasks: A survey, 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 486103d8-e4af-4e90-b429-11ed3f5d70a7 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Show, attend and tell: Neural image caption generation with visual attention, 2016
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 031a92e9-f738-4004-83a3-6d55f0356d40 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Lawrence Zitnick, Dhruv Batra, and Devi Parikh
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c85527a0-cded-4630-92f7-14800379efb6 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs GPT-4 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4658fd14-fb31-4d9e-aa57-4fc65c261901 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Qwen2.5-VL Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 577d6d7b-3301-454a-878e-baa80e0fd88d · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Deepseek- vl: Towards real-world vision-language understanding, 2024
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d5b3669-97ba-431a-81f4-3ecab4b991af · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0986aec-edef-4e42-b902-ccd81cca6c2c · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs 3D-GPT: Procedural 3D Modeling with Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a3d4937-1f3a-4212-bd29-b32096b4c599 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d678dc10-7f12-4c08-bc93-481545c78a48 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3113d389-c779-4717-8fdd-cb8e2be1fa06 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeca7fdd-8e26-4d22-9f5f-eec800e2f97d · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Detecting and mitigating hallucination in large vision language models via fine-grained ai feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 572ac910-67de-441d-8e6d-b4100e1d0a6d · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Rlthf: Targeted human feedback for llm alignment, 2025
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c44f92de-bc07-45fa-9df4-1788652c6064 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b8f7dad-9b9a-4e45-a218-35b8531d00d2 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Grpo-lead: A difficulty-aware reinforcement learning approach for concise mathematical reasoning in language models, 2025
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8011e1f4-495e-4e71-a275-5efd6e6df286 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfa8e2db-ff1b-47d4-84d4-4b85ce4f5afe · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Towards visual grounding: A survey, 2024
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64510dd2-04ee-4eac-9bc6-a3cffb2de1a9 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed228e8c-f02a-443c-a460-08c597959750 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Learning transferable visual models from natural language supervision
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a655876-6a4f-45e3-9a8d-dfafe75eea61 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da158765-d595-4c5f-9932-f501511c471a · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Introducing openai o1-preview
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 560d7927-4d9c-46b4-a9ea-3c2666b317b7 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47b2dc5f-db0e-4584-a974-4e11b75fa826 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs R1-v: Reinforcing super generalization ability in vision-language models with less than $3
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8ee8b4e6-4619-457f-98b9-377daf783e89 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79ac7303-9e85-429e-b7ef-91fbb3d54566 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e247608d-0ad2-45c4-a8a6-31f7482a5684 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Mm-eureka: Exploring the frontiers of multimodal reasoning with rule-based reinforcement learning, 2025
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ddce2ba-3ff0-4111-84da-ff185ab93f68 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Proximal Policy Optimization Algorithms
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b848829b-590c-4f91-a709-31e4c80891ab · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Referring expression comprehension: A survey of methods and datasets.IEEE Transactions on Multimedia, 23:4426–4440, 2020
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efc784ed-8649-4591-8297-215f26b94ec2 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs GREC: Generalized Referring Expression Comprehension
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f1f2259-9bc5-412c-86b1-6bbafb1b0725 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Vqa: Visual question answering
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93c76e6d-86e7-4cac-8e21-0a335cd68431 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Visual question answering using deep learning: A survey and performance analysis
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d1215f91-3097-4af8-a453-d810d496f0a3 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Generation and comprehension of unambiguous object descriptions
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbd0653d-6b7a-499c-b911-73639dfbb00b · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Modeling context in referring expressions
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b3c60dd-f435-4b9f-afd3-793571be997d · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Generating easy-to-understand referring expressions for target identifications
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c7e28c85-9174-4f25-9be6-d027b2389de7 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Lisa: Reasoning segmentation via large language model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83f6a806-eaf5-4f92-b3e0-654cc67cd64e · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs R1-onevision: A unified benchmark for vision-language reasoning and generation, June
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fb81a721-d494-473d-9364-7158b7b72664 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8df142c1-5998-48fd-ab18-c34ac98e8928 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs A diagram is worth a dozen images
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24dd8fe2-3296-47f5-833e-48517394d014 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Infographicvqa
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation df1babcd-f7df-45b2-baaf-0b279b825e02 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Towards vqa models that can read
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 513391da-8308-4512-95bd-4deba6eb9cb8 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Docvqa: A dataset for vqa on document images
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fba8c8d-445b-475a-b9a0-c406363c3b4b · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Measuring multimodal mathematical reasoning with math-vision dataset.Advances in Neural Information Processing Systems, 37:95095–95169, 2024
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21f31f13-f9d6-4de3-aa28-70385197870a · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Mmbench: Is your multi-modal model an all-around player?, 2024
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfa9161d-c6a9-40fe-813e-5ed6289c45ed · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs On the hidden mystery of ocr in large multimodal models.arXiv e-prints, pages arXiv–2305, 2023
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 50090ec8-945f-44e5-b00b-819a9d59a361 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation faf6aee4-a806-4df9-9f53-7c57f8dfc09c · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy, 2024
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 39b91a93-13e2-4b96-912d-4c0c9081f4f9 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e0354f2-5db9-4aff-9a28-3b87a75ddefa · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50630a9e-7574-4e54-9ed4-5cffbbbbb7f1 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 019551d9-c68a-4bf7-acad-7e23fd2cbdf4 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 090f1581-b4bf-455a-9812-f9122961da5c · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 930d8c78-b7c0-45b8-8292-63e104186493 · outbound
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5ae88b79-d5e3-4afc-bf32-456f4cc57230 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
Reference 249
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e9b9c12a-8212-4ce8-b8bb-139bacd918d8 · inbound
Faithful Mobile GUI Agents with Guided Advantage Estimator TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 643a9df2-e962-47c5-b9b4-fd16fbf27ffa · inbound
Trust Region On-Policy Distillation TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 20f70eeb-b24b-419c-91bd-eea014b96751 · inbound
Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fe678d8d-32ca-48e6-be9a-97308206e2ed · inbound
LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.