Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:27:27.827281Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 12 inbound Pith citation observations for arXiv:2505.04623.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:27:27.827281Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:13.645079Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T06:36:10.581673Z
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4b583074-11a8-439f-a37f-c14b260aa7d5 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Language Models are Few-Shot Learners
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2a31147-77fb-4652-b64d-67a13a4d6525 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52a38cfe-52a2-4db5-b38e-7d94f9226696 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Meerkat: Audio-visual large language model for grounding in space and time
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d0ff539c-5fd0-4e91-aaa5-997b077ec6c6 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 896af91e-17e4-4d06-9681-f14ead0631b9 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81be52e9-c157-40db-9323-3d4f55f1569b · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 604dcc0c-8dae-4bdc-acb1-9d2f44af38ba · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Aligned better, listen better for audio-visual large language models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f6786f8a-ce2f-41d6-9a5c-376375de6683 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3755e8f-7520-44c5-b5b9-398597adbc9c · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9deca145-4eb3-4eda-bfad-e25fc82b5eeb · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Reinforcement learning outperforms supervised fine-tuning: A case study on audio question an- swering, 2025
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3fd20983-66b3-40b1-9cba-1b76e40a456a · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning VideoChat: Chat-Centric Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ddb042e-a39a-467e-ba8e-c519944e414a · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning OmniBench: Towards the future of universal omni-language models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db271f2e-c079-482e-a40d-b05fe25f8bfd · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fa0bc64-62cf-4278-aafd-8ca2b1111978 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf64872e-1e45-4470-a8a4-3c4bd825e0d9 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55587098-4e55-4ebb-9224-1afc8103add8 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning ChatGPT: Optimizing language models for dia- logue
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 02d9c16e-2fd3-410e-85c1-8258fcfb539a · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Introducing OpenAI o1-preview, 2024
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation af057320-42e9-4ee4-b1d1-493c261e4fbd · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Training language models to follow instructions with human feedback
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 818984fb-e77a-47a7-a7f9-31f034ed8550 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Robust speech recognition via large-scale weak supervision
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 30d61480-77e3-40eb-a381-12f2b1ae17ec · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5f052796-5804-430d-a8eb-bd3a3f2c2dcb · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e45f8923-ab84-418f-b9ef-4a0d9f455b99 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29d67698-bd4b-47bf-9fdd-7277381470ef · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7c95f25-c2f4-4215-9faf-19278f05bbc6 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Audio-Visual LLM for Video Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 794e8067-f56a-4b2e-bde6-ff49c4c7946f · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Fine-grained Audio-Visual Joint Representations for Multimodal Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 963ef653-9f17-4b5f-a022-d3f4d8471171 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Hawk: Learning to understand open-world video anomalies
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c0f5dbda-e996-4b8e-9ae8-c033c9bea811 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Sari: Structured audio reasoning via curriculum-guided reinforcement learning, 2025
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2cec31b1-09f4-451e-99bf-792fb76c1a22 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c7eab77-d681-4a8d-a209-6c61a6932a84 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Audio-reasoner: Im- proving reasoning capability in large audio language models,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0e800d49-314d-4a38-8678-63d2554c9d14 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning EchoTraffic: Enhancing traffic anomaly understanding with audio-visual insights
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 14e74d12-223f-40c8-99e1-52830f61462b · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Qwen2.5-Omni Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fd210cb-f8a9-4e66-b6f0-a45c09f85238 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Avqa: A dataset for audio- visual question answering on videos
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a93e76ce-37e0-45b8-9e13-c96ca0fc2221 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bc1a75f-25e6-4a3b-b580-c93f4834e014 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a53c6b4-b4dd-47ae-b0eb-acb21f4e9081 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef2f2f44-9146-4cf3-ba64-eafcc5f6c11e · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7c1a011-9a13-42f2-a884-9b0061a923b5 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd67daff-5bc9-4a8a-8bf3-cf5a2b685f90 · outbound
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb8cd184-7364-4401-a920-23f1239dae01 · inbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8809531c-3f7a-4032-8af3-4ab46ef679fb · inbound
FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2caf0f52-2c2f-4cb8-bcf5-c365bbbc7f9b · inbound
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f29198a-1efa-4caf-8739-a82b313520b7 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
Reference 264
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e0a7d057-b8c4-4e62-9dde-6caf0ffe500f · inbound
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6cbf7b30-a75b-4436-97f2-cf6cb8b1392d · inbound
Development of a 3D-CNN-based Prediction Model for Migration Barriers in Plasma-Wall Interactions EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a9dfaff-cf0c-4434-995f-49833a1bc469 · inbound
Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cb2b7095-f2cf-4861-8605-52d8989c1bad · inbound
Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 87eebf59-fac5-45cc-b15c-4399c9236ab6 · inbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1afb3801-19f9-4cf1-a99e-336a9f694b0a · inbound
Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 43e2786f-801b-4e1b-aab3-a201e66cfbbc · inbound
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dcd6a649-c1a0-4046-af9f-6f01963efd5f · inbound
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.