Pith. sign in

Paper Citation Record · LEDGER

Ola: Pushing the Frontiers of Omni-Modal Language Model

As of 10 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 33 inbound Pith citation observations for arXiv:2502.04328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04328 v3

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:47:39.389717Z

measured 113 of 113 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:36:40.820699Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:57.266999Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36cc8bf1-f639-4d93-9b12-f766ef5737cc · outbound

This paper cites MusicLM: Generating Music From Text.

Ola: Pushing the Frontiers of Omni-Modal Language Model MusicLM: Generating Music From Text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.033803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.033803Z digest=sha256:d0fb89a07196112be450efc7b46c6219972e4cea2fc54caa5298fe4aa127c960

Observation 96c0d92a-1af9-4948-8e4b-d2c433f5928c · outbound

This paper cites Pixtral 12B.

Ola: Pushing the Frontiers of Omni-Modal Language Model Pixtral 12B

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.039633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.039633Z digest=sha256:11116b9a1796183986152b37b1cd3b7f684ec2161297755ed3af1d1901a83ef6

Observation e4eb3386-1687-4e99-a13f-255dd36cfdba · outbound

This paper cites Flamingo: a visual language model for few-shot learning.NeurIPS, 35: 23716–23736, 2022.

Ola: Pushing the Frontiers of Omni-Modal Language Model Flamingo: a visual language model for few-shot learning.NeurIPS, 35: 23716–23736, 2022

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.044339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.044339Z digest=sha256:ad2d7219c79d2193de1ed89af72f3a1c3b54263cdeec133c42db14b75842c00d

Observation c298e9bf-9dda-47d5-b750-d6aa875e20e2 · outbound

This paper cites X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages.

Ola: Pushing the Frontiers of Omni-Modal Language Model X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.048721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.048721Z digest=sha256:8af725c7700879cd90b03fdf235d1e20d47787cda3ec867c36e7aad530229c7f

Observation 339dcc89-68b2-4e8b-ab24-4856fa4e70ca · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Ola: Pushing the Frontiers of Omni-Modal Language Model GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.053243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.053243Z digest=sha256:c7eb56b1589a7023ba15f09ba3129d711f28b5602c4ef8e2a3ad53634de2cb14

Observation 1d05a51c-ac15-48bd-bfb8-255fc99a0ebd · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Ola: Pushing the Frontiers of Omni-Modal Language Model Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.057732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.057732Z digest=sha256:94a4573e54d962922abbe3bf82c3dec43495b367846bc98521045f109573eed9

Observation 84909031-5f6e-4f8d-8289-8991b2a172e7 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Ola: Pushing the Frontiers of Omni-Modal Language Model ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.062734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.062734Z digest=sha256:c99d261b42bdb059c890c161ceb617c6ebc607f8899ed3d3faaa909c8c7e0f85

Observation b69afcea-1c75-4a4f-8d2d-068d06171ca0 · outbound

This paper cites Beats: Audio pre- training with acoustic tokenizers.

Ola: Pushing the Frontiers of Omni-Modal Language Model Beats: Audio pre- training with acoustic tokenizers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.548206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.067887Z digest=sha256:eeac299865bd6e18b65e0877d1871df82e0b2237f938c9fe7fba0f7cf6556d0f

Observation d3d61bf3-5cef-4116-b793-1bcec186b665 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Ola: Pushing the Frontiers of Omni-Modal Language Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.073045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.073045Z digest=sha256:b61e57d39939b63d528b5777c7914ecad1f71b4da6b7d7f338894717ecfdc764

Observation aa5fd178-274f-479b-b61a-10d16037da86 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Ola: Pushing the Frontiers of Omni-Modal Language Model Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.533798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.077800Z digest=sha256:6788ad6392543b524c6ad9c220322f49dc8175f84e2706500aff325ace858f5f

Observation 0d6deb94-f2a1-45f5-a16b-c5b2bd9df8e6 · outbound

This paper cites Qwen2-Audio Technical Report.

Ola: Pushing the Frontiers of Omni-Modal Language Model Qwen2-Audio Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.082294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.082294Z digest=sha256:d8303cd5103d5a86903f99044fa2553485551134f87e17765c2f2fdccdfd52c3

Observation a8095656-7376-4254-90e7-aa126aa79d0f · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.087369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.087369Z digest=sha256:c093b520b7f3c5087ad008bcdd20839a28a62b299feba4d4a7fae10b3ad953ae

Observation be010433-2ee2-4dd2-8caf-8680846187d5 · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.092001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.092001Z digest=sha256:b1953577ebd091c5bc8041c629c8b8fad10cd9cc321dd3fe11f7fe2445cd84a3

Observation c4ac83af-cfa6-4765-bf09-a7c46045218f · outbound

This paper cites Clotho: An audio captioning dataset.

Ola: Pushing the Frontiers of Omni-Modal Language Model Clotho: An audio captioning dataset

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.519585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.096858Z digest=sha256:87aed06ffc3578a8f105ff2da8360058e8461740b88d705345e611ec20e5413a

Observation b2c86b57-3462-43f1-a363-34c80cbe1f48 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

Ola: Pushing the Frontiers of Omni-Modal Language Model Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.101465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.101465Z digest=sha256:45288492e818bb6310081383611cbe69c8868843dc7e7dc9f1cbfe093a1d7a91

Observation 7dc3286d-0748-4cc8-888e-601fe4d49c74 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.105733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.105733Z digest=sha256:2ba31353b9ac48685117a79c5817042831e92063a3f4ad43f238a27d5f48351c

Observation 9ec2ccbc-8a75-4cb3-9c5c-e8732075f457 · outbound

This paper cites Finevideo.

Ola: Pushing the Frontiers of Omni-Modal Language Model Finevideo

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.495231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.110748Z digest=sha256:5452bdb62ebb1ce2423e06fe2f906b33cb6622ad435423b4d64ad70dbfd9853c

Observation b6edf24a-58d0-40ea-83be-7dbb610912a8 · outbound

This paper cites Prompting large language models with speech recognition abilities.

Ola: Pushing the Frontiers of Omni-Modal Language Model Prompting large language models with speech recognition abilities

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.480758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.115161Z digest=sha256:35c2eb2b54ea3327a7e07cfd63d814264b894d65f6d247e516dddde3c0c725eb

Observation 748aef16-b096-4671-98f1-b15f73cda2a0 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Ola: Pushing the Frontiers of Omni-Modal Language Model Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.119516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.119516Z digest=sha256:aadab6cbf2540e21b466cabd710f2167005c48e2d7bde2afb88fc3f30339dcc8

Observation b20dffb3-3b4d-48fb-95ed-2a4a6244f084 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Ola: Pushing the Frontiers of Omni-Modal Language Model Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.124203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.124203Z digest=sha256:03d4686f1bbbe61f8da5351512177774974e7ee0e9159613cb700b6d724860e5

Observation 8cc93ee0-cec6-44eb-9cc5-923a6683089f · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

Ola: Pushing the Frontiers of Omni-Modal Language Model VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.128649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.128649Z digest=sha256:996a5821b064c9c6dbdfaf322c8c050ee760a2bd742b2753bf3cbee545ede67a

Observation 284e8d0c-f632-4704-9b80-307f553280a2 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Ola: Pushing the Frontiers of Omni-Modal Language Model VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.133157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.133157Z digest=sha256:1e9162c5687b523893efae9b32f7315911a0a664ebec8b613570073394e4813c

Observation 9190eed9-e584-49e2-b687-6d0b3c4bd1f7 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Ola: Pushing the Frontiers of Omni-Modal Language Model Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.138130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.138130Z digest=sha256:bae9a45d6b4541d8d625694094b5c71808c183ed3c4faaf6411506396977bd86

Observation 32799e57-afca-408e-97c5-16b96a1b73a1 · outbound

This paper cites Hallusionbench: an advanced diag- nostic suite for entangled language hallucination and visual illusion in large vision-language models.

Ola: Pushing the Frontiers of Omni-Modal Language Model Hallusionbench: an advanced diag- nostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.142824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.142824Z digest=sha256:f83f6c58fc89bb6a0b27b56a42e258bbdf5e337c067b92f1278573fe6685f117

Observation f02c1b66-f08c-445d-a481-2881b210d20a · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

Ola: Pushing the Frontiers of Omni-Modal Language Model MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.147822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.147822Z digest=sha256:c3190340656637c96d020a44b1e18b7e549c5249dc706369f8f7242f03731baa

Observation 68d3c0cc-b976-4704-b1ed-ae04ece4bc1e · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.NeurIPS, 36:20482– 20494, 2023.

Ola: Pushing the Frontiers of Omni-Modal Language Model 3d-llm: Injecting the 3d world into large language models.NeurIPS, 36:20482– 20494, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.456893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.152244Z digest=sha256:3976ecffdff0d451ce76f6b5c58fe65f55996343e0997c000555403ca820a5a1

Observation 3286a960-6612-4779-9a7d-7be08835858f · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head.

Ola: Pushing the Frontiers of Omni-Modal Language Model Audiogpt: Understanding and generating speech, music, sound, and talking head

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.442436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.156237Z digest=sha256:bf185bcb2eaecb508d3795fd1b40380aa96e3be2628ddd8a66a20288277a8364

Observation 915809bb-967a-49ec-9b7b-587b34b1c388 · outbound

This paper cites A diagram is worth a dozen images.

Ola: Pushing the Frontiers of Omni-Modal Language Model A diagram is worth a dozen images

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.428020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.160152Z digest=sha256:691f73a1d1f9753b563edf2f7e645038341fb5891d5274dfb2c3c9a3f0a85cc2

Observation cd595f54-6dfd-477a-be1f-5417575e5ac8 · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

Ola: Pushing the Frontiers of Omni-Modal Language Model Audiocaps: Generating captions for audios in the wild

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.413912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.164136Z digest=sha256:91931ca84aabea613df5f192e74bfe5c39602370c0756d2d6e38f5eca9736663

Observation 094c9579-4d98-4f7c-ac56-95046e30840c · outbound

This paper cites What matters when building vision-language models?,.

Ola: Pushing the Frontiers of Omni-Modal Language Model What matters when building vision-language models?,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.168130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.168130Z digest=sha256:9a202304f5d274f7972c557876e879336401983b035d3f1a2ef09801103d832a

Observation 2ccad04d-6854-46c4-9f16-3d646fad8786 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Ola: Pushing the Frontiers of Omni-Modal Language Model LLaVA-OneVision: Easy Visual Task Transfer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.172390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.172390Z digest=sha256:3fd52cf30b0900cd44d44c7caeab76659da182489729cfbd2ed88f025c11ece0

Observation c7a3cc94-b33c-4581-ab86-4b75029c9ab9 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Ola: Pushing the Frontiers of Omni-Modal Language Model Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.390244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.176335Z digest=sha256:687470cf7012337b5a54f1e58fdd5eda57f7684f2891ddfe5ab71a9ed724cb46

Observation e2c42e2c-b04c-4a94-8f70-42b3e1f71714 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.180712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.180712Z digest=sha256:b4f5242dfb1127a1d8e4f53b77cf2b666217a2867bed3b53521d6136566251b9

Observation aa6285d0-f875-4fac-b0b9-9ac1cdc58209 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Ola: Pushing the Frontiers of Omni-Modal Language Model Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.185155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.185155Z digest=sha256:5257812823d68ec52697a0bc412d8a31516a6e0d0e0c909cd90fee02162f0004

Observation e1c9e7d7-3850-434b-b9d2-bda94cdf279d · outbound

This paper cites Improved baselines with visual instruction tuning.

Ola: Pushing the Frontiers of Omni-Modal Language Model Improved baselines with visual instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.376428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.189638Z digest=sha256:a6b8086b0c866d221b23452bc23827ed55db2bd03d218f25bcd058d76fd77abe

Observation 42a524a5-c017-416e-b7d9-00b89c4ce79c · outbound

This paper cites Visual instruction tuning.NeurIPS, 36, 2024.

Ola: Pushing the Frontiers of Omni-Modal Language Model Visual instruction tuning.NeurIPS, 36, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.362427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.193766Z digest=sha256:99e51ab1e278d198b2a7e3f26e1191a97d8008c5b405e621bf2147939a932142

Observation 6a7e792b-09d0-461e-8e87-0be6e8c963b9 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Ola: Pushing the Frontiers of Omni-Modal Language Model MMBench: Is Your Multi-modal Model an All-around Player?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.197952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.197952Z digest=sha256:13441404b448399dd0065007ccbc550be423832154534b9128876c4e2d7ff4a1

Observation 3f1ea753-738a-42b7-9dd4-3a7ad0fe66d2 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.202509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.202509Z digest=sha256:a2cd1da4416b024fcec096dbcffcc5ea85bfe00be648b52d8c8ae35262c0f200

Observation 0d28cb45-166f-4728-9025-ba01d2f44225 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Ola: Pushing the Frontiers of Omni-Modal Language Model Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.207697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.207697Z digest=sha256:25df4044108800de992a11cd4bba405fb936d80c89d3a7133c7063008e301d5b

Observation 12ae62c6-8899-485a-86ff-e0ce6deaca68 · outbound

This paper cites Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.212186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.212186Z digest=sha256:62e940bc6c730c0ed1a40625c29681aca10ef6ea0a82c4f1c533b1e778e5b00d

Observation 2ef09257-3459-4409-ae06-ddae9877efd8 · outbound

This paper cites Efficient Inference of Vision Instruction-Following Models with Elastic Cache.

Ola: Pushing the Frontiers of Omni-Modal Language Model Efficient Inference of Vision Instruction-Following Models with Elastic Cache

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-08T22:47:39.762001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.216914Z digest=sha256:194148f1caeb6de2f1a1cad3cfb6f9e56a599ffa1dab84ca0c7f9a09c56dba81

Observation 285adbf9-99ec-436d-88a5-531da218938b · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Ola: Pushing the Frontiers of Omni-Modal Language Model DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.221310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.221310Z digest=sha256:9c87ddb1c4b1f90ccf074e74df792b233028f19b7202eaec4fbc43bd3e08a11f

Observation 7bbe32ce-e38a-4d57-af59-e3947f23078f · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Ola: Pushing the Frontiers of Omni-Modal Language Model MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.225870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.225870Z digest=sha256:4610770c22bf906dc0a87e4a820b6c2eb6d9d68e80c232253ed163f4ca023d47

Observation 73c510a9-e742-4f9e-b557-48e913ef9a77 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.230344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.230344Z digest=sha256:d2a0bf3db6855722e79f5255af04b6ac9f7ad9464b205b37b8bba98ac81d819f

Observation c7000482-891b-43e8-9b24-80a146338ddc · outbound

This paper cites The million song dataset challenge.

Ola: Pushing the Frontiers of Omni-Modal Language Model The million song dataset challenge

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.348049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.234810Z digest=sha256:ffc8039464483e95c80f6eb8cb922299e6d56cf2bb9efd3ade9ae283e73d3081

Observation d33f4b2f-646c-4215-9fd7-429c655d951f · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multi- modal research.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024.

Ola: Pushing the Frontiers of Omni-Modal Language Model Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multi- modal research.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.333894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.239231Z digest=sha256:cfdee3452110f358c8e3b98048a4ab71e02383f02057e4754e690dcaa39a7845

Observation 468d8b70-41a2-4d10-9a9e-cf29947a2a4f · outbound

This paper cites Openai gpt-3.5 api.OpenAI API, 2023.

Ola: Pushing the Frontiers of Omni-Modal Language Model Openai gpt-3.5 api.OpenAI API, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.319172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.243492Z digest=sha256:1b86110f55c3b3981d8620f1657c74ce302da448df6bca290a4776cdbda7f062

Observation f5ac6ca2-80f7-4bab-b80c-c1ac36009b3b · outbound

This paper cites Gpt-4v(ision) system card.OpenAI Blog, 2023.

Ola: Pushing the Frontiers of Omni-Modal Language Model Gpt-4v(ision) system card.OpenAI Blog, 2023

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.305194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.247874Z digest=sha256:553e6407af6033511fbb7f948bc27c6dd0b1e188e3d40584a6cada2d8fde769e

Observation b14035cc-a05b-41da-aa4a-11279cf1e9eb · outbound

This paper cites GPT-4 Technical Report.

Ola: Pushing the Frontiers of Omni-Modal Language Model GPT-4 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.252511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.252511Z digest=sha256:c75076f0466ebb1474714597357d08262794f538ec149e1870e4118ff814f008

Observation f2a69b28-10c8-4593-8d99-0192b7704f5c · outbound

This paper cites Hello gpt-4o — openai.OpenAI Blog, 2024.

Ola: Pushing the Frontiers of Omni-Modal Language Model Hello gpt-4o — openai.OpenAI Blog, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.289737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.257343Z digest=sha256:9a06e48cf22cca15405515d38fc22c134a807cd2e9a6d7668a19f32d200a0bd0

Observation 2d958a66-1976-49d6-a5c8-c63f236b158d · outbound

This paper cites Librispeech: an asr corpus based on public do- main audio books.

Ola: Pushing the Frontiers of Omni-Modal Language Model Librispeech: an asr corpus based on public do- main audio books

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.273985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.261725Z digest=sha256:2431d8bb8fb034bec1072eb383e02c827de932c6b61bf89a49d34a25ce2d6972

Observation ebde3cab-2cd6-4a00-bb8a-fb189d31bb3e · outbound

This paper cites Streaming Long Video Understanding with Large Language Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model Streaming Long Video Understanding with Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.266066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.266066Z digest=sha256:d2ced0b4a425d9fcce64209b91670cffc728c6d568e3c3a351ea7cb04195200f

Observation e6e01e05-6c44-437d-b3ea-3e9c669289ca · outbound

This paper cites Qwen2 Technical Report.

Ola: Pushing the Frontiers of Omni-Modal Language Model Qwen2 Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.270605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.270605Z digest=sha256:ff2edfb1f28c167ea473496d8e6d0487e904692dec6d973527c9eb3c4dd85508

Observation dbf09d75-de8a-41a8-a55e-6ba4d5f982f0 · outbound

This paper cites Qwen2-vl: To see the world more clearly.Wwen Blog, 2024.

Ola: Pushing the Frontiers of Omni-Modal Language Model Qwen2-vl: To see the world more clearly.Wwen Blog, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.259014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.274734Z digest=sha256:467aa17066ffa98f5d5a138a3f7eb185f163e9ee9af2bdf9ba5a18e029f887cd

Observation 5135a1fb-ca51-431e-90c1-90af14ff915c · outbound

This paper cites Robust speech recog- nition via large-scale weak supervision, 2022.

Ola: Pushing the Frontiers of Omni-Modal Language Model Robust speech recog- nition via large-scale weak supervision, 2022

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.243530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.278590Z digest=sha256:78bfed87aca5242595424f2d079a51abb3a7a6dc59f2ccb1f57f63c4c085c55d

Observation 6bc40731-e56a-4617-ab74-3b1a1ff2a2bd · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

Ola: Pushing the Frontiers of Omni-Modal Language Model Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.228446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.282567Z digest=sha256:114f1aa502727bfe1738c02da3ab515e63c2a2228c6fa67a8e0d00fae2d33e01

Observation ba1cb269-c287-4983-b8a5-afd2497a9b36 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

Ola: Pushing the Frontiers of Omni-Modal Language Model CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.286819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.286819Z digest=sha256:f7f28196f35a7c4efaedbf043191024cac7a67fcfea54f0470f871a9a58596f6

Observation 6a7c2c05-3021-4866-989e-068aa7196854 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

Ola: Pushing the Frontiers of Omni-Modal Language Model AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.290979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.290979Z digest=sha256:b47a2d407dfda42d0e3aa632cdd01080293bb7f73c431f1d112d4c53e9f8390b

Observation 92bffe21-3310-495a-91bf-cb568b7dbb45 · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

Ola: Pushing the Frontiers of Omni-Modal Language Model MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.295397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.295397Z digest=sha256:4f75a3d083efe0d89f0c7e086ba46b0f4c9f26a3d41e42ca908b1a4b282f4bab

Observation f4e70091-ef1b-4385-8c25-f8fa97886be9 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Ola: Pushing the Frontiers of Omni-Modal Language Model LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.299864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.299864Z digest=sha256:cc315dda84e1e414c1ed6a6e0cba347c5edaa34c33705ab064499d0bf914119e

Observation b62ffc24-67a2-48c9-8536-aabd360bf92c · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.304398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.304398Z digest=sha256:47a33020f87967e76dc4a16f37a23c6e941d5cfdee485d953f737f160ba9ff16

Observation 8f29f496-73dd-4629-a25a-71cfb294925d · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

Ola: Pushing the Frontiers of Omni-Modal Language Model Qwen2.5: A party of foundation models, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.213910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.308845Z digest=sha256:df1110a2fdb88fadfb1545a677e6ec8de10a810a92d3d5652f04386d8e2a0a74

Observation f02690d6-cb74-4241-99e1-3fd1f20758c7 · outbound

This paper cites Qwen2.5-vl, 2025.

Ola: Pushing the Frontiers of Omni-Modal Language Model Qwen2.5-vl, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.312940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.312940Z digest=sha256:dc4ea6bff3deb94b2fa826543fe2b3384621d5a678ce8bd8c7298fdfd054c860

Observation 9655bb93-6ecd-4ba6-a0df-4dc8aec363b0 · outbound

This paper cites Learning Features of Music from Scratch.

Ola: Pushing the Frontiers of Omni-Modal Language Model Learning Features of Music from Scratch

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.317222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.317222Z digest=sha256:be6a4c14a2ccceaca8a791ce559ee1020bec343f94b5472a5ad8dad3d3d7cb5e

Observation dcf6dc59-a89d-43bf-9aa5-a1cf5dd525de · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Ola: Pushing the Frontiers of Omni-Modal Language Model Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.322045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.322045Z digest=sha256:1c22de2c4d2950065a570fe949bfa0ee3fe75c31cc9fa8cfc08b79e0c7c2f01c

Observation 21dda9e6-26c1-4c46-ab0d-76e8d652c64f · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Ola: Pushing the Frontiers of Omni-Modal Language Model LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.326631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.326631Z digest=sha256:fa756b440a5d0522e0e3f515210249a0deabdcb03a03802f5ba4279a3dfcf4a9

Observation b5dbb05a-f09a-4431-9b3f-2c114c7bc473 · outbound

This paper cites On decoder-only architecture for speech-to-text and large language model integration.

Ola: Pushing the Frontiers of Omni-Modal Language Model On decoder-only architecture for speech-to-text and large language model integration

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.186955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.331408Z digest=sha256:ccf71d027fe89d92cc791ca5761c2533953689f0c9eee9b8f5b7a8bcd35386ca

Observation e74f88dd-b47d-41e1-bc18-36eb5ea138e0 · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

Ola: Pushing the Frontiers of Omni-Modal Language Model Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.335713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.335713Z digest=sha256:4759201b5fd2818541d5ea7d10fbc5cb39496b401081adfdb2a80635b44ffed0

Observation b9c0e938-8242-4b36-a5e8-c13da6e66357 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Ola: Pushing the Frontiers of Omni-Modal Language Model Qwen2.5-Omni Technical Report

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.340402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.340402Z digest=sha256:00470b51a3fca09053afc20b06b0e74b23a6af4d272117536383279f28d98e16

Observation 8d754839-e1c2-444b-b95f-24a45f959baa · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

Ola: Pushing the Frontiers of Omni-Modal Language Model AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.344979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.344979Z digest=sha256:1eeaa464e43bbcc370e1b1eca8b2a3e996e050e2ad7b8d20009ccce4e45d5c27

Observation 1e41759d-1964-4f45-a720-32c7192f653a · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Ola: Pushing the Frontiers of Omni-Modal Language Model MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.349659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.349659Z digest=sha256:d124eae215c7278ae9a647ce9f702ad1858d505ae016ba041c98348c39827cd4

Observation 4fa3a959-f049-4aa8-8f20-891d7b801c66 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Ola: Pushing the Frontiers of Omni-Modal Language Model Yi: Open Foundation Models by 01.AI

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.354193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.354193Z digest=sha256:252d84689f8297c0dc83d06121c99aa432f17d0c86319f8cc0f9a97a99d1e130

Observation 8b956f44-8267-4cab-af4c-f303836b5dbb · outbound

This paper cites Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi.

Ola: Pushing the Frontiers of Omni-Modal Language Model Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.172036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.358992Z digest=sha256:a4eee2ccdbf3cdecaa50cf4a94d87c78d95d16142a83dc0e82d3c238d8404059

Observation 4024c18d-223e-408b-bbd8-b86158815af2 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Ola: Pushing the Frontiers of Omni-Modal Language Model LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.363415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.363415Z digest=sha256:d6ff51e3c61ff01e74f050e7e0ba65cb64204b0cb4796cc1b7d7dfd09f078dd3

Observation 15082f6f-e3ea-42c3-b2b2-4273c8c5f284 · outbound

This paper cites Sigmoid loss for language image pre-training.

Ola: Pushing the Frontiers of Omni-Modal Language Model Sigmoid loss for language image pre-training

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.156938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.368080Z digest=sha256:dd8841669589e1f32ff617cfb6321dc6abba3f297baa218b19dcc29cb59c798b

Observation a5d9ac97-e599-4e74-ab34-e903bb88d8a9 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

Ola: Pushing the Frontiers of Omni-Modal Language Model SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.372296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.372296Z digest=sha256:0dcefbd5dcdccf7e23b5f855f4324c06968f2b53a98c7ae796f2dbf3a77a3cb6

Observation c4b67586-0fda-4dfd-a13e-849b0c11f5c8 · outbound

This paper cites InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions.

Ola: Pushing the Frontiers of Omni-Modal Language Model InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.376946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.376946Z digest=sha256:4942cddc0df8fea426f7381cca193356a7fe23adaf65efd23bb844ab4aa5df26

Observation 005b69e8-1a87-462e-a221-1724ddd9ccd9 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

Ola: Pushing the Frontiers of Omni-Modal Language Model Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.381215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.381215Z digest=sha256:454f5fc6fb6458b24b3e11a61aa03903e028c981a8cbc4d7df39f01f0990b9ff

Observation e7907262-478e-44d1-b798-74a35c42a3f6 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

Ola: Pushing the Frontiers of Omni-Modal Language Model Llava- next: A strong zero-shot video understanding model, 2024

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.142589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.385607Z digest=sha256:2fc20ed7ca4498c2794335040f20724ab77053d6f82943015f73b88950f69ae4

Observation 77d877b2-5f39-43b3-9717-b46f71729b2a · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

Ola: Pushing the Frontiers of Omni-Modal Language Model Video instruction tuning with synthetic data, 2024

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.126837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T22:47:39.389717Z digest=sha256:489f296214d1a46065591c692d9b3aa398c86998920b93c0d524805f2368bf9f

Pith citing papers

Observation 72593941-1480-4b2f-8cae-5df4ae47d81a · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:34:36.856851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:c3a2a73db86241101221945b9ad4210a95cba9f1e984571ebec6c27f4cd41dc2

Observation f4be1992-5bfe-4944-9db6-92afa7afd6d1 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T01:00:51.428555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:eb0fc85a6251c7a588791f23d280a68d8786608c7420fd9e6ea0f304ce4b6350

Observation 79da6943-9a46-463a-9bf6-86ecd9b7955a · inbound

Is Extending Modality The Right Path Towards Omni-Modality? cites this paper.

Is Extending Modality The Right Path Towards Omni-Modality? Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:40.820699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:40.820699Z digest=sha256:ba54ae548a84c07d2db4dd75edc5deaa2a7825c4c2dc991126429a28865159ae

Observation be396407-bbd2-4d64-bb62-6ae72aab8163 · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.533023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.533023Z digest=sha256:9880ca3437c592cf8def0fa997a27ef3bb24ac0c0d09b834c5bd259c6bb7ae6d

Observation bf4ec20d-c45d-4fb2-90a6-085f09707196 · inbound

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs cites this paper.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.562512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.562512Z digest=sha256:9ddaa55ea81463acc831b777306aedee17163122dbc9201c1ee81bd48db1af32

Observation 757897fa-3c40-4072-bca9-bb4530bc2848 · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.719441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.719441Z digest=sha256:ad65817e23a8fd58e69753a46975915a49617a51bc51086f6e5127d3f521968c

Observation 6a229ab0-3baa-4764-bed8-d65f74e078af · inbound

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning cites this paper.

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:32.454824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:32.454824Z digest=sha256:f2871fa5aebb1793236f1f208165d60492212eea1874f73e0079d6b21ae7915a

Observation b21fc665-8563-4dac-84d1-1b4f15f5c7a2 · inbound

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context cites this paper.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:13.250680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:13.250680Z digest=sha256:4e1ed92973f3aff19de17070c4a09ccb113c6fdb5611d8a8223f59e5cec7cb24

Observation 891797d4-fed9-483a-87f6-fb8f24729b7f · inbound

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning cites this paper.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T06:12:07.011738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:628be725d8561951c5b24471e08f906a4663fb2dfcd969f6de1757429fa6bf23

Observation 37c6c68c-dfa0-458c-aa6f-d438fb959ff9 · inbound

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding cites this paper.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.446078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.446078Z digest=sha256:3dc4ce74bdaf7d15fac5ee78aa33cc46f0471b1ec4943dc08bc598b9d17dd612

Observation 96ea2f47-d8aa-495b-8e66-fb699670c22b · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:07.832466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:07.832466Z digest=sha256:b6ba051e0993eca0105f010b3f4cc98ec79c965784ff1c302371625f55d1d6a8

Observation 93a144e0-8404-42a2-b380-4c3252808c8b · inbound

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning cites this paper.

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:21.059602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:42:21.059602Z digest=sha256:57474550201329cff421467da530beb4f8ae7ad9969760ffa837c8136cc4a300

Observation 1b28e4ed-9481-4ad9-9b68-ffef3d04fb0f · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:40.028245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:40.028245Z digest=sha256:2a0da9d45022fa0ccfb98ca7835f8cf614d8b4c2e15579f6fab223e00adfaeaf

Observation 52c170a9-6577-4502-9730-ac4b1ecd25c6 · inbound

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs cites this paper.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:01.361938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:01.361938Z digest=sha256:296ad6fc6429d7c5ff5dc8b3edf7e589f43c104b1cb5e5260fadeb45a92f68bf

Observation feb94f13-ae33-40bc-a428-480a624c18e5 · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.170122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:8e93a6636be18494cb0cb819a47e5c81b474b81a2acc677bb60a3463aaf18535

Observation 2c6d5228-e79f-41a6-952e-3860e9100e21 · inbound

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers cites this paper.

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:18:32.227284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:02:53.574120Z digest=sha256:1a7060a1131a5215897cb4b100814b07f728afc4a885a25c9d36ee5482b5b0bc

Observation 23330239-3260-46ab-bcb0-7c3a5a358b14 · inbound

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective cites this paper.

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 281

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:51:03.196437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T04:32:29.428080Z digest=sha256:c08ea75eb22c1b9fa3b39cf42474fc379c106639a1a32176d8f65ae863fb440a

Observation 1f256362-ffdb-4a6b-bfac-9af589abdda7 · inbound

Valley3: Scaling Omni Foundation Models for E-commerce cites this paper.

Valley3: Scaling Omni Foundation Models for E-commerce Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:51:05.976588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T14:53:55.160230Z digest=sha256:f3aa22a869d6ccc212151daf8766382268dd3b63b6b7999987fdeb7fdb7b3747

Observation bea1c93c-b9e3-4145-bf8c-0242fa2c665f · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:15:55.936975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:15330cbdcdaf253066e3f7198e7886e3846ceb5ae54a9024aac68cb965e430b6

Observation 576aaaa5-9692-4b5a-bcf5-e9d11540dc1b · inbound

Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs cites this paper.

Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:07:33.673297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T18:06:27.962891Z digest=sha256:5cfe17182b9752f85358628de5b1c4488fe5f6b39320352b21e65fd52e87ee30

Observation e310f04a-6116-4a6b-a3ad-cf50d7278c41 · inbound

Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing cites this paper.

Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:53:24.776257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:53:11.761843Z digest=sha256:6c7b99ba41c869ce5880bcd5ed9c6ef3c7f785141e51fa2990c2480aa38ab626

Observation e046fc4b-1d3e-439a-8325-8d14bb80a7b7 · inbound

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding cites this paper.

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:55:24.901284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:51:49.390597Z digest=sha256:fe8f21528664b7aad28ef9914ee1ee1cb15ae1066fee48770b7f8ef51c7a8242

Observation 3843e84b-bbca-4776-8646-153ed1d55dc8 · inbound

Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain cites this paper.

Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:23:17.992169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T10:20:12.860927Z digest=sha256:40a85cb693dc014332a3c5751cbb8b9c00ac7e929a277ac21ab1a59ef28fc7b6

Observation 1ed0ee24-4c98-4e50-97f0-f325d22baf69 · inbound

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation cites this paper.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.177260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:a32dbb08543a2177e56a5f0a79305d43fe9c46028247ef1d4eace31cbc610566

Observation 65eca983-78dc-43d8-af42-ee4c9ea5a895 · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.391014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:c84162aeeace3c33b0ff36aefddf16fe98b38a16d4c9fec978257c23ddc34b53

Observation 0a3d6fd9-e5aa-4f55-851a-05d55517200a · inbound

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs cites this paper.

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.365632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T16:12:53.387567Z digest=sha256:ab77beed8b609e7abfa1f12c84439fb1f58c478a207198815673fd42850e3129

Observation ba0ac130-32b2-431e-a8a1-0f689da9c32a · inbound

MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning cites this paper.

MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:58:03.433035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:43:56.299302Z digest=sha256:dd1e2ed6138bff8a1c840e57ebd7863f0189d8701f7986fb373428eb3c2634c2

Observation 7c30ffa1-b114-4e1e-a1ce-e3adfb50705d · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.386746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:a3beec22510e5a8194289788ce05029f100e0a7e7a045d57b54efac33605698a

Observation ae2729d0-56ea-4491-963f-4bd682948047 · inbound

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression cites this paper.

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:49:57.269489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T00:16:56.174638Z digest=sha256:d009c5745d3d0de91820b54f7e75cc8aeb897d97a792ae8913c1d961a23ac3b9

Observation 53305da3-9d94-4f90-9e0b-b2b6ea9cba5f · inbound

RedVox: Safety and Fairness Gaps in Speech Models Across Languages cites this paper.

RedVox: Safety and Fairness Gaps in Speech Models Across Languages Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:52.872965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T04:37:00.399470Z digest=sha256:cf1e5e8dad964379cce511e9b31662af5172cb367bce388799a5ca416968b8a8

Observation 367c2df0-13f8-499a-b7d2-0f7a77a9917a · inbound

Conversational Human Audio-visual Talking Dialogue Generation cites this paper.

Conversational Human Audio-visual Talking Dialogue Generation Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T06:59:35.258176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:59:35.258176Z digest=sha256:2328dacb5a88332d209662a81d4bcb059de7c4afdd14f4d29ca277ed1b82570b

Observation b27d6c40-0dd5-4988-85f7-3028d1c79fa7 · inbound

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning cites this paper.

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T05:48:27.255331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:48:27.255331Z digest=sha256:570d3c829a46d68ec98c7e96f6abaf5532fe4e458f336d1dc8ad32e90315a1a5

Observation 66981f95-7734-4775-aed4-d246dfee7d91 · inbound

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos cites this paper.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:50.700260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:50.700260Z digest=sha256:1121a597eadfcfadac6f68b679e9df7619ad2abc7727961e55fab73fcbfafdf3