Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:41:12.567839Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 4 inbound Pith citation observations for arXiv:2506.20960.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:41:12.567839Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T03:12:38.454127Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-28T17:42:26.143560Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f225cebc-1161-439d-8716-a5d59df990eb · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e2366f1-1b8a-42b4-9d74-84e83b6181e8 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Lawrence Zitnick, and Devi Parikh
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d587128-15c0-4c15-810d-3476be033fec · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs From pre-training to fine-tuning: An in-depth analysis of large language models in the biomedical domain
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd096af0-f108-46e6-9c71-25011fc1ec36 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline, 2017
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0c1fecd-0d5e-46e2-96fd-1bf79aa69b5e · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beb48705-2f65-4847-94ff-a778753b0bd1 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98545770-40e3-4b0b-bc3c-91bc9998f5e0 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs EmotionLines: An Emotion Corpus of Multi-Party Conversations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b275f25c-94bb-4ada-ab5e-4bd3a85d5e37 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a666ecfd-3861-4c73-84bc-c7c195e80402 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1b67d40-cf4b-446a-8d72-87c15ab0214b · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Training Verifiers to Solve Math Word Problems
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 683b0b58-bb8b-431d-a76c-cfc12a45be5f · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Fleurs: Few-shot learning evaluation of universal representations of speech, 2022
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ab4286a-3fb9-438e-aa33-a05606cd76c0 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Mmbench-video: A long-form multi-shot benchmark for holistic video understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a5efb7c-427e-43b3-8172-acffd2c5344a · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Finevideo
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7ebc80b-8bb9-418e-8453-3897aa102aad · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92ec66cc-80e7-40c1-ab3b-b1c9d3761695 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fdc8698-1508-428c-81bc-49d4698ad336 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a23a837b-7bfe-465a-8445-9b69c65e6321 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Gemini 2.5: Our most intelligent ai model, 2025
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c1095ca1-58ba-4ecf-b3c5-920ab8acce7a · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a4c75c9-2cc5-478f-98a8-8d6243394041 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs VizWiz Grand Challenge: Answering Visual Questions from Blind People
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caf43fd7-4a25-488e-aff0-47b4d871774b · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Measuring Massive Multitask Language Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24130295-32dc-4437-8bd4-1a9d53dd9372 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs TED-LIUM 3: Twice as Much Data and Corpus Repartition for Experiments on Speaker Adaptation, page 198–208
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a45e1a72-0485-486f-a6ef-fbca14729fec · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Worldsense: Evaluating real-world omnimodal understanding for multimodal llms, 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86971764-831a-49ba-b2e2-67cb9fdfa03c · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1067f61e-484e-4a6e-acb5-5e50ea6be5fc · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Weld, and Luke Zettlemoyer
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91af4e98-0bde-44e4-9fb7-3cec99c1f56d · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a62604c-a5a0-4205-8515-e12c0ed20d7c · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36e74076-1271-4dd2-98aa-4daa1a72c325 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Baichuan-omni-1.5 technical report
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b5ba30b-c55a-4e90-8ea9-5408d8b54426 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Evaluating object hallucination in large vision-language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cdc5fcfe-06de-4bef-99d7-b1027923e643 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Omnibench: Towards the future of universal omni-language models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de84453d-89d0-429f-be37-bec0b7aab529 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Omnibench: Towards the future of universal omni-language models, 2025
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06bc4f0d-0227-4fc5-9358-d442d0b8ea74 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 885f63c8-9367-4978-9fb4-894fe5dcf70f · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Clotho-aqa: A crowdsourced dataset for audio question answering, 2022
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6be41f8-518f-4f04-ab54-ebb24811f76f · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Visual Instruction Tuning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb6f2337-76e9-4cfa-a6fc-2fc12be03d4a · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs MMBench: Is Your Multi-modal Model an All-around Player?
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c640d218-a471-4267-a56e-a0243a26a596 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b404fbb-252f-44f4-b006-15416d8bec1b · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 57c96686-2be7-4790-88ff-73314df941a9 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Spoken question answering and speech continuation using spectrogram-powered LLM
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 633ee0bf-d760-4d6b-be59-1f7fdf72682b · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs VoxCeleb: a large-scale speaker identification dataset
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70251e21-663e-4876-8a5a-e6596219200f · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Minicpm-o 2.6: A gpt-4o-level mllm for vision, speech, and multi- modal live streaming on your phone
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01e21f72-be9b-4566-8e57-04744ff39443 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Librispeech: An asr corpus based on public domain audio books
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88e76809-59f1-413b-accd-badad0c908f7 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Plummer, Liwei Wang, Christopher M
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f400eb72-c47b-4d0c-a95c-99896722aad0 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs MELD: A multimodal multi-party dataset for emotion recognition in conversation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 572dbd7c-2733-4873-b333-0f651e326e04 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Question-Answering Dense Video Events
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59b980dd-af92-4650-873a-384245114850 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Learning transferable visual models from natural language supervision
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48d88fd4-04d5-4c15-8189-57fc55a9ba9d · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Towards VQA Models That Can Read
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9b2e1ea-2cde-4190-aaaa-c0fc00120c60 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs A precise detection method for transient micro short-circuit faults of lithium-ion batteries through signal processing
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a9893c17-2fa5-4e67-ae51-9b576e6ccc2c · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Language Models are Few-Shot Learners
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f235de8-3e1a-48f2-994d-218f17a1f298 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs GPT-4 Technical Report
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b785eea0-6ead-4bd2-9624-86372ae71d2a · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Qwen2.5 Technical Report
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c07cb19b-8a30-4d36-a3b7-d64e99640254 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs CogVLM: Visual Expert for Pretrained Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90f6c13b-1feb-4c9b-9e49-f33c3b804798 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50918665-cf97-4c2f-9f92-a8529f81cb66 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40f2cc99-5c71-49af-843d-2f38b6f1c043 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Qwen2.5-Omni Technical Report
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01de9829-d349-4f13-972c-fb4f35bb5ecf · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Air-bench: Benchmarking large audio-language models via generative comprehension, 2024
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b3e6c2e-49c4-4f9b-8419-3f622fc5d2e1 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.TACL, 2:67–78, 2014
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0bd4c52-9490-47df-89e2-74c141147d0c · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Berg, and Yuandong Tian
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6222991a-cc70-4303-8a1e-e0bbd2d9d79c · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f47c004c-2c97-4e69-bf3a-a2ee632fd0f3 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition, 2022
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21460dda-f689-4f1d-940c-7ec25356ea79 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Zhang and M
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 524268b1-1fcb-4478-a98d-3003119ad138 · outbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Yin and Yang: Balancing and answering binary visual questions
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2730f9b8-4354-44bb-b738-6016ba2b65de · inbound
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db3424f7-a96b-489d-bee4-18522e05b557 · inbound
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5e4067b-2c44-4678-a1c2-06322a667f58 · inbound
Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb12b0a5-33cd-4e82-98f0-de806c00ccc2 · inbound
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.