Pith. sign in

Paper Citation Record · LEDGER

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

As of 7 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 4 inbound Pith citation observations for arXiv:2506.20960.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20960 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:41:12.567839Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T03:12:38.454127Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T17:42:26.143560Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f225cebc-1161-439d-8716-a5d59df990eb · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:06.734883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:06.734883Z digest=sha256:1b33e88375570f63cd24cce699aaad72bc411dd6ef19baa8a7cce52c4ee0e148

Observation 5e2366f1-1b8a-42b4-9d74-84e83b6181e8 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Lawrence Zitnick, and Devi Parikh

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:16.535302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:06.815550Z digest=sha256:dc5de8879b13923a1c2a31d90cb5bcaee3632bcdbf986dab4cb337d95fcd3a83

Observation 1d587128-15c0-4c15-810d-3476be033fec · outbound

This paper cites From pre-training to fine-tuning: An in-depth analysis of large language models in the biomedical domain.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs From pre-training to fine-tuning: An in-depth analysis of large language models in the biomedical domain

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:16.366029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:06.894106Z digest=sha256:4a352766f3cfeb85ca9cf810487ac90257d4763250e2707deb2b3c43fcffee95

Observation cd096af0-f108-46e6-9c71-25011fc1ec36 · outbound

This paper cites Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline, 2017.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline, 2017

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:06.985286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:06.985286Z digest=sha256:1edf0e4e32b72bb5aa8b4a6bb61db583ff2a960c7dc07abcd9a74daa1f6ebddb

Observation f0c1fecd-0d5e-46e2-96fd-1bf79aa69b5e · outbound

This paper cites CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.070388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.070388Z digest=sha256:a7f18590cd69e1ad0019313e1c87c483bd45dee6361073b272b338eb8d16d9c5

Observation beb48705-2f65-4847-94ff-a778753b0bd1 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.185729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.185729Z digest=sha256:f024d3bf0bff1506d277f2f2c6d747f7d2747fb9284df52fb8b6214cfa6b22cf

Observation 98545770-40e3-4b0b-bc3c-91bc9998f5e0 · outbound

This paper cites EmotionLines: An Emotion Corpus of Multi-Party Conversations.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs EmotionLines: An Emotion Corpus of Multi-Party Conversations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.259689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.259689Z digest=sha256:8f3b750f8277b468969f5594cdbed4da6813042289bfa6d6f5f406a0091db897

Observation b275f25c-94bb-4ada-ab5e-4bd3a85d5e37 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.346165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.346165Z digest=sha256:9fc9a168c4072beae08ad4d33265dc0573bc6a31891ef4778cf0eba82629f516

Observation a666ecfd-3861-4c73-84bc-c7c195e80402 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.424327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.424327Z digest=sha256:7d72cde202e805cae63be9706aa10d26bc2f4728385b029d968fa32869da6bef

Observation d1b67d40-cf4b-446a-8d72-87c15ab0214b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.513119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.513119Z digest=sha256:2fb0df937bef5f8aa850aba859d0d66effc8834cf4cf8129823414ba02a31d42

Observation 683b0b58-bb8b-431d-a76c-cfc12a45be5f · outbound

This paper cites Fleurs: Few-shot learning evaluation of universal representations of speech, 2022.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Fleurs: Few-shot learning evaluation of universal representations of speech, 2022

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:16.225901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:07.564476Z digest=sha256:be63c7b81539eae5c0b15b9091a496ee73d6923546b74f204a50179382a6f8dd

Observation 4ab4286a-3fb9-438e-aa33-a05606cd76c0 · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Mmbench-video: A long-form multi-shot benchmark for holistic video understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.676280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.676280Z digest=sha256:a31609a3c0be06ff4d084fe5ffdcd04d512d1be360c3f630132b6da82f29a21b

Observation 3a5efb7c-427e-43b3-8172-acffd2c5344a · outbound

This paper cites Finevideo.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Finevideo

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:16.029585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:07.761731Z digest=sha256:ef2a263a7a227aad884cda9dd74c4db3d98405ee1dcd69f9c22b20fb5631a37c

Observation f7ebc80b-8bb9-418e-8453-3897aa102aad · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.861174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.861174Z digest=sha256:664c773eef63d02f5c93b7a53ddfc76057d9abaf94d01ae91e7b7cb767ad4697

Observation 92ec66cc-80e7-40c1-ab3b-b1c9d3761695 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.943880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.943880Z digest=sha256:3445faf0f87b7d2faf70ddf8b6b8137ec128e328ab6234862cd9359e17e5e433

Observation 7fdc8698-1508-428c-81bc-49d4698ad336 · outbound

This paper cites LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:08.047046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:08.047046Z digest=sha256:fa4b0254e8c6e13d5b3b0c5a3e21fd13b1c99c26a3f91c42c654f502013d64fa

Observation a23a837b-7bfe-465a-8445-9b69c65e6321 · outbound

This paper cites Gemini 2.5: Our most intelligent ai model, 2025.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Gemini 2.5: Our most intelligent ai model, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:15.895029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:08.151171Z digest=sha256:a57402e4ec92757c1209ece5eeaef43b67ffb4c3fa1556a514312d9db9a9aa6b

Observation c1095ca1-58ba-4ecf-b3c5-920ab8acce7a · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:08.251757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:08.251757Z digest=sha256:093b2b6964ed7ef4a8424f1c9ba0d2002794cfb6f32b82fa06c8ee8f54c1e027

Observation 4a4c75c9-2cc5-478f-98a8-8d6243394041 · outbound

This paper cites VizWiz Grand Challenge: Answering Visual Questions from Blind People.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs VizWiz Grand Challenge: Answering Visual Questions from Blind People

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:08.363669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:08.363669Z digest=sha256:f6f0ee434917b928d2860f39a95bd10045fef0172d4d7992a0064489a260d18d

Observation caf43fd7-4a25-488e-aff0-47b4d871774b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Measuring Massive Multitask Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:08.456014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:08.456014Z digest=sha256:09b481241edfff37460a03450893b3e59c18854ae95e317b321a44480aadd264

Observation 24130295-32dc-4437-8bd4-1a9d53dd9372 · outbound

This paper cites TED-LIUM 3: Twice as Much Data and Corpus Repartition for Experiments on Speaker Adaptation, page 198–208.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs TED-LIUM 3: Twice as Much Data and Corpus Repartition for Experiments on Speaker Adaptation, page 198–208

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:15.690697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:08.535756Z digest=sha256:032649c7080596b728be35df928fce7d7ccb7e890a9ba3f32ab26c07f112c578

Observation a45e1a72-0485-486f-a6ef-fbca14729fec · outbound

This paper cites Worldsense: Evaluating real-world omnimodal understanding for multimodal llms, 2025.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Worldsense: Evaluating real-world omnimodal understanding for multimodal llms, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:08.604727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:08.604727Z digest=sha256:b9692cf87c1f875f0d4f34b18e74a29ff61c88482c979940d16822d9b4dc9938

Observation 86971764-831a-49ba-b2e2-67cb9fdfa03c · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:08.706606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:08.706606Z digest=sha256:5c898532ee6b4c6025de365ab229aedc1a6d271c2889991c1bb42c129a673a4e

Observation 1067f61e-484e-4a6e-acb5-5e50ea6be5fc · outbound

This paper cites Weld, and Luke Zettlemoyer.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Weld, and Luke Zettlemoyer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:08.818096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:08.818096Z digest=sha256:208492985cd3a9b082d1db1dcb8078cf972155d53d4ac70b623d081319bc70dd

Observation 91af4e98-0bde-44e4-9fb7-3cec99c1f56d · outbound

This paper cites an unresolved cited work.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:41:15.536276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:08.896095Z digest=sha256:7492f31432870e5c98d82041ec3448c479c072640c6040b78aed64ece96cf569

Observation 7a62604c-a5a0-4205-8515-e12c0ed20d7c · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:15.399615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:09.002699Z digest=sha256:29d452cdb5eade474f4aeb3ac460e617fd7f2e30c42e08fbfe21c8944cb7ed7f

Observation 36e74076-1271-4dd2-98aa-4daa1a72c325 · outbound

This paper cites Baichuan-omni-1.5 technical report.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Baichuan-omni-1.5 technical report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:09.124008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:09.124008Z digest=sha256:e962a5c17570e5fe333db756754c59ae0769b5b68ea8ec7ce76c3fdce1a01bf4

Observation 3b5ba30b-c55a-4e90-8ea9-5408d8b54426 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Evaluating object hallucination in large vision-language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:15.251245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:09.236197Z digest=sha256:1a346a2dd9f5953b65053396770f9efec85ffe46b24e10fd17445f031a928b7e

Observation cdc5fcfe-06de-4bef-99d7-b1027923e643 · outbound

This paper cites Omnibench: Towards the future of universal omni-language models.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Omnibench: Towards the future of universal omni-language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:09.320715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:09.320715Z digest=sha256:edf55df842c95ec47929474b45fe37a3e4849efd232dcf17ed4a61696d9303b5

Observation de84453d-89d0-429f-be37-bec0b7aab529 · outbound

This paper cites Omnibench: Towards the future of universal omni-language models, 2025.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Omnibench: Towards the future of universal omni-language models, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:09.397134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:09.397134Z digest=sha256:e93ebf94ab0a7e02d1ba4642cfc84089f2ccdc8c8cf99fedb378dd74faa407d1

Observation 06bc4f0d-0227-4fc5-9358-d442d0b8ea74 · outbound

This paper cites StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:09.479989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:09.479989Z digest=sha256:de2828ed43f4201256d648e0432c2c491bb87dc3aad90a33eadf29373a168834

Observation 885f63c8-9367-4978-9fb4-894fe5dcf70f · outbound

This paper cites Clotho-aqa: A crowdsourced dataset for audio question answering, 2022.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Clotho-aqa: A crowdsourced dataset for audio question answering, 2022

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:15.107148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:09.570093Z digest=sha256:888db57b8b8d4ea0c5ec996e11a713bddae62f173be8d71063a8570761b7652e

Observation e6be41f8-518f-4f04-ab54-ebb24811f76f · outbound

This paper cites Visual Instruction Tuning.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Visual Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:09.665514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:09.665514Z digest=sha256:02a3e4bfbe69f9394764517f8d255dcfb7e7475862e24832201ca0e98ad7011f

Observation bb6f2337-76e9-4cfa-a6fc-2fc12be03d4a · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs MMBench: Is Your Multi-modal Model an All-around Player?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:09.774966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:09.774966Z digest=sha256:80f7b7b99263aba8313678449a43d4fe71bad615cc1dded039831342c2fcc894

Observation c640d218-a471-4267-a56e-a0243a26a596 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:09.864310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:09.864310Z digest=sha256:1d1fda4a2aba5fa3cd263a203d027e38dc09f99feb93396782e909075d5b524b

Observation 8b404fbb-252f-44f4-b006-15416d8bec1b · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:14.948481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:09.951816Z digest=sha256:6c08ba4735aad54f6e0d4998c608fc22bbde808ca87655bbd9a0cd7ab513c962

Observation 57c96686-2be7-4790-88ff-73314df941a9 · outbound

This paper cites Spoken question answering and speech continuation using spectrogram-powered LLM.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Spoken question answering and speech continuation using spectrogram-powered LLM

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:14.746493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:10.082574Z digest=sha256:e632e08131e0c0ee4a49d2540aa4aebd2465543f86fbad6e2c290c53262cee68

Observation 633ee0bf-d760-4d6b-be59-1f7fdf72682b · outbound

This paper cites VoxCeleb: a large-scale speaker identification dataset.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs VoxCeleb: a large-scale speaker identification dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:10.216134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:10.216134Z digest=sha256:a6b754dd059b9be1aa69173021f44eba7aea67c4839b3b2e8d62da6a8099ebff

Observation 70251e21-663e-4876-8a5a-e6596219200f · outbound

This paper cites Minicpm-o 2.6: A gpt-4o-level mllm for vision, speech, and multi- modal live streaming on your phone.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Minicpm-o 2.6: A gpt-4o-level mllm for vision, speech, and multi- modal live streaming on your phone

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:14.604431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:10.316996Z digest=sha256:5e139e92919902bc92e523aaa7a1b42547bc4cf1f53f011c2490e7851072120a

Observation 01e21f72-be9b-4566-8e57-04744ff39443 · outbound

This paper cites Librispeech: An asr corpus based on public domain audio books.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Librispeech: An asr corpus based on public domain audio books

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:14.410005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:10.447338Z digest=sha256:86eacbaae1e70eebf4cb2355cd7d33e994487047486337c364286522800d465a

Observation 88e76809-59f1-413b-accd-badad0c908f7 · outbound

This paper cites Plummer, Liwei Wang, Christopher M.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Plummer, Liwei Wang, Christopher M

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:14.264172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:10.542832Z digest=sha256:d59ae722d47db692e2f044ca1b79547a8e654472bc54a5f9e31c426726f88ac9

Observation f400eb72-c47b-4d0c-a95c-99896722aad0 · outbound

This paper cites MELD: A multimodal multi-party dataset for emotion recognition in conversation.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs MELD: A multimodal multi-party dataset for emotion recognition in conversation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:14.121328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:10.673146Z digest=sha256:11dd44ed26496570e47d5469258d10963c6ba9225883895033c8f125d7c10314

Observation 572dbd7c-2733-4873-b333-0f651e326e04 · outbound

This paper cites Question-Answering Dense Video Events.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Question-Answering Dense Video Events

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:10.769779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:10.769779Z digest=sha256:a2327bb7219c71ec864ce8cab2af5a49967ba4280f7afa895d3ab190399649c1

Observation 59b980dd-af92-4650-873a-384245114850 · outbound

This paper cites Learning transferable visual models from natural language supervision.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Learning transferable visual models from natural language supervision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:10.938665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:10.938665Z digest=sha256:b44ed35063e157c9ca26ab5be2d994f70bf3cf84ffd5b8fab215ad2fcb8bce82

Observation 48d88fd4-04d5-4c15-8189-57fc55a9ba9d · outbound

This paper cites Towards VQA Models That Can Read.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Towards VQA Models That Can Read

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.063492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.063492Z digest=sha256:42ad19f722fe43d213d4c3957dc39e2fc1fff162fd853934761b2557cce8cbea

Observation b9b2e1ea-2cde-4190-aaaa-c0fc00120c60 · outbound

This paper cites A precise detection method for transient micro short-circuit faults of lithium-ion batteries through signal processing.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs A precise detection method for transient micro short-circuit faults of lithium-ion batteries through signal processing

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:41:12.954686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:11.176834Z digest=sha256:e1e5ec15c34e0a3e17de0a68fb5b48eeb787817b7c9dae5c92c746bbb84f6167

Observation a9893c17-2fa5-4e67-ae51-9b576e6ccc2c · outbound

This paper cites Language Models are Few-Shot Learners.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Language Models are Few-Shot Learners

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.279409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.279409Z digest=sha256:b03e5909b7bddd9b1dc3edb2296677e9b554fe0cea0c5314077b821b7187231a

Observation 7f235de8-3e1a-48f2-994d-218f17a1f298 · outbound

This paper cites GPT-4 Technical Report.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs GPT-4 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.369391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.369391Z digest=sha256:ec08a8ed84bce9a4b39d5be340c33dd7178d2391c09434cec7883e63aded2afe

Observation b785eea0-6ead-4bd2-9624-86372ae71d2a · outbound

This paper cites Qwen2.5 Technical Report.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.475765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.475765Z digest=sha256:6b12c28005ef3efcc7a02f2ccf883c256d08cbe38b62d160fc5c2f447c9b6d54

Observation c07cb19b-8a30-4d36-a3b7-d64e99640254 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs CogVLM: Visual Expert for Pretrained Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.561442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.561442Z digest=sha256:b7bc8e159715291108dfb81c44c567a85649e72c282ca6acd8d7a7f1f77155ad

Observation 90f6c13b-1feb-4c9b-9e49-f33c3b804798 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.644932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.644932Z digest=sha256:e7f266b0a068333cc6779cf1d968c2e71d455eb3a7450f9c01ceef71ccc5aebd

Observation 50918665-cf97-4c2f-9f92-a8529f81cb66 · outbound

This paper cites Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.661900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.661900Z digest=sha256:711b2e766012e6e1009a244d2f06850fef1c2afa8b345a7de09ceaa1d0880815

Observation 40f2cc99-5c71-49af-843d-2f38b6f1c043 · outbound

This paper cites Qwen2.5-Omni Technical Report.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Qwen2.5-Omni Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.841114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.841114Z digest=sha256:c5f9f056513b18dbb3057d72bf2ac2c1829468cb3785d13af6675b73d0781c99

Observation 01de9829-d349-4f13-972c-fb4f35bb5ecf · outbound

This paper cites Air-bench: Benchmarking large audio-language models via generative comprehension, 2024.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Air-bench: Benchmarking large audio-language models via generative comprehension, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.967533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.967533Z digest=sha256:35c47be82b57b5e11d3d9c0f48ded5f81478fb614f5028964718c570ec141e04

Observation 2b3e6c2e-49c4-4f9b-8419-3f622fc5d2e1 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.TACL, 2:67–78, 2014.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.TACL, 2:67–78, 2014

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:12.118730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:12.118730Z digest=sha256:916dd25936550b3ad87655beb3bbb156cbdf5f84c7f787ed349a416305684493

Observation d0bd4c52-9490-47df-89e2-74c141147d0c · outbound

This paper cites Berg, and Yuandong Tian.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Berg, and Yuandong Tian

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:13.941233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:12.231890Z digest=sha256:55a478dbafaeb6fe351278b42da4a8249194d970c358eee2c28a39a38f45e9e0

Observation 6222991a-cc70-4303-8a1e-e0bbd2d9d79c · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:12.314340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:12.314340Z digest=sha256:eebc2131d0984b726c76b40821570c63e48fb95b04e084e3a5ae472a23086b96

Observation f47c004c-2c97-4e69-bf3a-a2ee632fd0f3 · outbound

This paper cites Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition, 2022.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition, 2022

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:12.402858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:12.402858Z digest=sha256:e3eef848c1deb07edff972646997685fafa8ef1e5e373661ed64640a7ff2640c

Observation 21460dda-f689-4f1d-940c-7ec25356ea79 · outbound

This paper cites Zhang and M.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Zhang and M

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:13.810971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:12.455460Z digest=sha256:2ba716ecdf966c7d8d8a58e1f6f418b9717b09ffcdf5759e34dcaae118322b73

Observation 524268b1-1fcb-4478-a98d-3003119ad138 · outbound

This paper cites Yin and Yang: Balancing and answering binary visual questions.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Yin and Yang: Balancing and answering binary visual questions

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:13.679803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:12.567839Z digest=sha256:14f47a62575eb0d61e8931aa15152346a3d743f424bbb9a43a801d3138e9535b

Pith citing papers

Observation 2730f9b8-4354-44bb-b738-6016ba2b65de · inbound

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models cites this paper.

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:58.971096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:58.971096Z digest=sha256:682d6bc751f161b3a47c56234c1a6801a194417c3911ebd9c03d05806713cb91

Observation db3424f7-a96b-489d-bee4-18522e05b557 · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:56.078287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:99d10404a5dd28db65a2773191e6c3d247036e15daf2358836b4109395be7265

Observation c5e4067b-2c44-4678-a1c2-06322a667f58 · inbound

Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition cites this paper.

Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T17:42:26.145281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:40:30.038337Z digest=sha256:cc94421ad1a2c8a79f4e594a424082a8a465ae945543b45b2cab7545f76d2d1a

Observation cb12b0a5-33cd-4e82-98f0-de806c00ccc2 · inbound

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation cites this paper.

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T03:12:38.454127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:12:38.454127Z digest=sha256:98b2bc1dd29544e9002b19d55226c8d03e491f66570d98a859517b69fd7d6733