Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2311.18799.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:48:15.981135Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:09:41.279523Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 6f02d1db-a0d8-40b6-b7b2-d537c641e7b9 · inbound
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2017dba7-a8e0-4ae1-85a3-df5a1fdfb783 · inbound
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7d9c0ca-fa84-4fcc-9c2c-24f50449ee61 · inbound
Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af7ff0c7-ffa8-4aa4-919c-5a1ac56f26e4 · inbound
Modality-Inconsistent Continual Learning of Multimodal Large Language Models X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2c48730d-aea6-4791-95e7-6cf2672d3188 · inbound
A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 280
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ee2955f-56a1-4c70-a346-44c9a35a680e · inbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54232287-b629-49b0-9e68-f757043a9a7e · inbound
3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3f1422e-6964-464e-9a20-bf4909b6da1c · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 160
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2089806d-d4d8-4faf-9925-14246f765018 · inbound
Foundational Models for 3D Point Clouds: A Survey and Outlook X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 176
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfebbb1a-da40-42bf-b345-d6b3fb941bfb · inbound
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a17bc5c5-59a7-4933-b169-1d602e105c20 · inbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b9bed74-cfd5-43d2-81ed-4a1fdc659e68 · inbound
AuthGuard: Generalizable Deepfake Detection via Language Guidance X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e11c5f6-141d-4a64-b401-f753d12fb7db · inbound
Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 191b7aef-7ad7-4a45-8d33-809fc4b8aa07 · inbound
SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74b74614-b132-4d43-ad93-f59553198c90 · inbound
CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bee4125a-b824-4a19-a701-e96c733a7aa2 · inbound
Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fdfb7a6-5a81-47c0-a436-b1e74a46124e · inbound
Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation acb3a1a4-d757-4d89-a3fd-b84f2b0363c0 · inbound
Closed-Form Spectral Regularization for Multi-Task Model Merging X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 04d39992-225f-4a31-be9b-a27bbdc21052 · inbound
CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.