Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:33:19.591854Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 5 inbound Pith citation observations for arXiv:2506.13642.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:33:19.591854Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T06:35:35.951554Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T16:38:39.602552Z
79 of 79 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 99f2d984-0e85-404f-af46-d47bad206dc3 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Hello gpt-4o, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4ca0a7a8-44db-47f7-ba61-10986922dad3 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Gpt-4v(ision) system card, 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cb87d4dc-dd24-4c03-98d4-d07b0c69ad42 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Visual instruction tuning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 491b9874-b198-45ae-8e7b-ab812f1ef63d · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model MiniGPT-4: Enhancing vision-language understanding with advanced large language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5c78f174-2af0-4e89-8c5c-5555031f56af · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9c087cb-4c3b-42f8-9a46-81de19f1e5a2 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaVA-OneVision: Easy Visual Task Transfer
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ab3b4fb-bae9-4a7b-a7d3-1e62ea7e4305 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1b7e80d-0a19-4bc4-9c34-91e24f446e9a · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7448f4ea-2857-4560-beb2-e3769993f9bd · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaMA- omni: Seamless speech interaction with large language models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e1373ac1-5099-4818-8963-e2ba1604d6b3 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Moshi: a speech-text foundation model for real-time dialogue
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 171ac415-b9c8-47de-aab3-1bd0b16a5aec · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c0249fb-3f19-4f57-bc5e-ef16e778e3a3 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00c9d13f-5346-4bb2-ad89-999189f95508 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Baichuan-omni technical report, 2024
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0adcf620-700a-46e1-a44e-26ca606ba35a · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Qwen2.5-Omni Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88f62f36-3081-4824-8222-4e2ba58ee3fe · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Connectionist tem- poral classification: labelling unsegmented sequence data with recurrent neural networks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77b2176b-144e-4162-aa55-adcd1168d97c · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Learning transferable visual models from natural language supervision
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e545fb63-b739-4379-885b-1e155e47775a · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 162c306d-abff-46dd-a029-c6edbdfa953e · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 16c684df-b327-4f82-8bb9-7e17c22b4528 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2828e8b0-ff6c-4883-b663-885ba785378b · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 90ff68e9-ab1c-4fff-90a3-784606db2304 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd4ad7ce-5871-44c9-9b2f-0bc008ee470e · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model VideoChat: Chat-Centric Video Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aea73c1b-d24c-42f7-9447-41c286b980ea · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Video-ChatGPT: Towards detailed video understanding via large vision and language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fa43ef72-8084-4d98-b812-6fccee3b5343 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d03987c5-bb12-41b9-896b-94b8880e4c21 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Video-LLaV A: Learning united visual representation by alignment before projection
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e40c7e43-5d2a-4814-a945-19f891f788fb · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09ba4a5e-31cf-4a00-89cd-cc489715c8a5 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a75a3f1e-d172-4cc1-a04d-c8d055653d95 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Slam-omni: Timbre-controllable voice interaction system with single-stage training,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7e87629b-a43e-499e-b0e3-1baac53bd90e · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Robust speech recognition via large-scale weak supervision, 2022
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 20ba2d0d-0fac-499b-a766-14073c675d5b · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76a7d486-85e6-46e4-a154-e4a7113a5e5d · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Hubert: Self-supervised speech representation learning by masked prediction of hidden units.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3451–3460, 2021
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 364e34da-2e2e-4130-ae6e-9ecf434d071f · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40a80281-7f94-4d0c-a57c-851ee32bd298 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1393f3cc-df50-425f-a3c8-4714f50a3c8c · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Speechtokenizer: Uni- fied speech tokenizer for speech language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 49e0837c-b280-4bda-9abb-ed83e6a41216 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b572f75-01bb-4bdd-8a94-1835bf8faf8e · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6907a3a9-4e83-4aa4-a288-2d6d4560034d · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79c360b9-12cc-47f9-8e9d-f0c67c6d7099 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Megrez-omni technical report, 2025
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3f5707ab-84f4-4a61-aa79-9f94835d864c · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49d2bc70-bf19-47eb-9c8f-432f1887f8fb · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5f8eeb56-c015-44ca-b3ee-76677e515931 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2337fda3-cc8c-4241-b296-8134f263461a · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Openomni: Advancing open-source omnimodal large language models with progressive multimodal alignment and real-time self-aware emotional speech synthesis, 2025
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eef4a63-d46b-45e7-b69a-d26e6226ab37 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Stream- Speech: Simultaneous speech-to-speech translation with multi-task learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aec48e8a-01c8-456d-98c8-0c62539af473 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaV A-mini: Efficient image and video large multimodal models with one vision token
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7b3b7fc3-c782-4dd2-9b12-c42d5f2977ea · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Future-guided incremental transformer for si- multaneous translation.Proceedings of the AAAI Conference on Artificial Intelligence, 35 (16):14428–14436, May 2021
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0d4c36c6-dd8c-4162-a85d-e910384fbe28 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model STACL: Simultaneous translation with implicit anticipation and controllable latency using prefix-to-prefix framework
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b886f9d2-8782-415d-98a7-ca81cdb87e1f · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Information-transport-based policy for simultaneous trans- lation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed288ccc-6968-47fe-9e67-c6d5fdd2ceaa · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Universal simultaneous machine translation with mixture-of- experts wait-k policy
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 030d498b-c61a-4151-8daf-91be67b987b6 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Librispeech: An asr corpus based on public domain audio books
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cad4eaf6-cc97-4b89-922c-a8f33a02d995 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model End-to-end simultaneous speech translation with differentiable segmentation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9ddbca23-aa03-4346-b39c-71a5b8142f41 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8b54d6d9-9aaf-4136-a244-4645ab435966 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47529764-308a-4bec-98c9-0cf8a7c46da4 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 15dbb48e-f8d2-4abf-a228-c45ec74ff014 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Hudson and Christopher D
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b4715a3b-9719-4bfc-8ba2-92d0d69d98ad · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Towards vqa models that can read
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ce3a6731-d55c-4499-acf2-7f72fe21f344 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Learn to explain: Multimodal rea- soning via thought chains for science question answering
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9ab882af-001f-4b7d-9ed5-96fe76b00cae · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8475a16-8b23-45cb-a4c0-978f7933fbb5 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Evaluating object hallucination in large vision-language models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6f30e2b9-f600-4094-b497-c145b6453040 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Seed-bench: Benchmarking multimodal large language models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a170ddd9-daa1-4f37-8cb4-3e6b391b00ef · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model MMBench: Is Your Multi-modal Model an All-around Player?
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e0579f0-f8fe-4228-ad87-283cd2471329 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Mm-vet: Evaluating large multimodal models for integrated capabilities,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 783463fc-6429-43de-aa46-6cdf7ab10579 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Visual instruction tuning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a224011-d38c-4c0b-8e36-661891981140 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Semantic parsing on Freebase from question-answer pairs
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3af33f61-c3d4-48a9-bffb-067b3f7454fe · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Visit-bench: A dynamic benchmark for evaluating instruction-following vision-and-language models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f5487bc2-ab3e-4bc9-b9d6-0d435ab04a66 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Spoken question answering and speech continuation using spectrogram-powered LLM
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cf89375b-0d49-4cba-8248-1ed91245a399 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Improved baselines with visual instruction tuning
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40133eb9-41ff-4f21-979f-69e2e2a4f8ce · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Introducing idefics: An open reproduction of state-of-the-art visual language model, 2023.URL https://huggingface
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f2e0cdb4-aef5-43c3-a9d0-eae9f60b6d8e · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Textually pretrained speech language models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation abe7db92-287e-4465-9595-8a98817694c7 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cfecd7a-cb94-4f44-ba50-8e016548ba32 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model The Llama 3 Herd of Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a119a8b9-f13b-4860-8699-792aabae8383 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Sigmoid loss for language image pre-training
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5e67540-2a28-4247-9b5c-61fabd5ceea5 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 387faf8f-1424-435a-b860-dc09ef65783b · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b913fec9-ae1a-4c18-bfd8-c3f2b8284f66 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Enhancing chat language models by scaling high-quality instructional conversations
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22b2724e-132d-4889-9947-a15400715054 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model does not allow traveling to the second floor
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d2c1a34-e4c4-48c0-b0cf-f4ee13d27e4e · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Scaling Speech-Text Pre-training with Synthetic Interleaved Data
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c062e03-716e-412a-8997-ad76e25128a1 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model URL https://aclanthology.org/ D13-1160/
Reference 2013
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8ede076c-5894-4de9-8ed9-890ce438eb60 · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cba64abf-8252-40a7-9f0b-0e0ebac5daca · outbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model doi: 10.18653/v1/2024.emnlp-main.342
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 286ba852-3dfa-43a3-9a01-4abb28a5dfab · inbound
See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 20d11b5b-424b-468b-98b5-45b0698792ac · inbound
PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b3a8cf48-e086-43d3-bc8e-60ddc7414b0b · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 95af9d12-ec95-40ab-8db1-5aad1f844823 · inbound
Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f28a4588-daff-4cac-8c3e-8020f4f57764 · inbound
Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.