Pith. sign in

Paper Citation Record · LEDGER

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

As of 18 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 5 inbound Pith citation observations for arXiv:2506.13642.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13642 v2

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:33:19.591854Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T06:35:35.951554Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:38:39.602552Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact3
  • verified fuzzy30
  • unresolved44
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99f2d984-0e85-404f-af46-d47bad206dc3 · outbound

This paper cites Hello gpt-4o, 2024.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Hello gpt-4o, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.616625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.211972Z digest=sha256:9290da524f1c8825f98863406539fa3e4e2c49ff4e88331d73589e2dd4a59f64

Observation 4ca0a7a8-44db-47f7-ba61-10986922dad3 · outbound

This paper cites Gpt-4v(ision) system card, 2024.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Gpt-4v(ision) system card, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.579132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.217907Z digest=sha256:8561d30d597b2a30151108737c94b3daab5c0a7a09e9e60867deefa48e7b17c6

Observation cb87d4dc-dd24-4c03-98d4-d07b0c69ad42 · outbound

This paper cites Visual instruction tuning.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Visual instruction tuning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.534744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.222673Z digest=sha256:8549ea97a75516cd4ecddb15118b4fac3195993db9059ca2bbece4cfefe0acdf

Observation 491b9874-b198-45ae-8e7b-ab812f1ef63d · outbound

This paper cites MiniGPT-4: Enhancing vision-language understanding with advanced large language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model MiniGPT-4: Enhancing vision-language understanding with advanced large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.490643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.227705Z digest=sha256:812f17a84574522883aa6912e6297b9b2858d0f805927b3b269eda49a0108c4b

Observation 5c78f174-2af0-4e89-8c5c-5555031f56af · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.232753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.232753Z digest=sha256:a7b4e018930d265899c9da3e71f0c8776458e98991507b499293429ee17866b4

Observation b9c087cb-4c3b-42f8-9a46-81de19f1e5a2 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaVA-OneVision: Easy Visual Task Transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.237792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.237792Z digest=sha256:8df389a095261699d97d35c3c3deb1dae5f9cdde07a69918bb4293e094416388

Observation 1ab3b4fb-bae9-4a7b-a7d3-1e62ea7e4305 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.243615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.243615Z digest=sha256:2c1db5dddadb349083fcdc788f35e446fd9bba873bc11a91851d2cdc3129babc

Observation c1b7e80d-0a19-4bc4-9c34-91e24f446e9a · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.248709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.248709Z digest=sha256:e0bef39930a104d7a5fae5fe723091c39d1d61117bd0660e714f01b2bc4121fe

Observation 7448f4ea-2857-4560-beb2-e3769993f9bd · outbound

This paper cites LLaMA- omni: Seamless speech interaction with large language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaMA- omni: Seamless speech interaction with large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.452144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.253606Z digest=sha256:ca3c074c8071acdf604f7ee6fb3a11d640c1c87dfc6629c59dba50254ff01be3

Observation e1373ac1-5099-4818-8963-e2ba1604d6b3 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Moshi: a speech-text foundation model for real-time dialogue

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.258298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.258298Z digest=sha256:0953c337282c7aef93d582389d6d0b8c5f7c5308ac97c9c9f26d7ccc80c3adcf

Observation 171ac415-b9c8-47de-aab3-1bd0b16a5aec · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.263029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.263029Z digest=sha256:166a490e286dbdc3c6d2fd29962d22fc1d12fadd2fb4c201638ae846bc7556e1

Observation 3c0249fb-3f19-4f57-bc5e-ef16e778e3a3 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.267903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.267903Z digest=sha256:7d6832af65a7b60069eccfe255adce6914f5cc13c4a213ed7ea1b011589049f9

Observation 00c9d13f-5346-4bb2-ad89-999189f95508 · outbound

This paper cites Baichuan-omni technical report, 2024.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Baichuan-omni technical report, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.273419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.273419Z digest=sha256:8d0d4b9cba49b3495fcfb64ce5562c40a332cbb6a4fd2b28b533c1c01a3746eb

Observation 0adcf620-700a-46e1-a44e-26ca606ba35a · outbound

This paper cites Qwen2.5-Omni Technical Report.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Qwen2.5-Omni Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.277910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.277910Z digest=sha256:4dfa531784c949e1bee4cb370ced272660261a7e7614135f2ee55e7189344dd9

Observation 88f62f36-3081-4824-8222-4e2ba58ee3fe · outbound

This paper cites Connectionist tem- poral classification: labelling unsegmented sequence data with recurrent neural networks.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Connectionist tem- poral classification: labelling unsegmented sequence data with recurrent neural networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.282753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.282753Z digest=sha256:530cb4456c2858f4b03c83f44d820804339b49417741994ad2a8798c01ea82e7

Observation 77b2176b-144e-4162-aa55-adcd1168d97c · outbound

This paper cites Learning transferable visual models from natural language supervision.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Learning transferable visual models from natural language supervision

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.431137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.287814Z digest=sha256:8cc0fe5b47440f4339a090c432b04d72657bc2822c3c8841553e6d785f676e3e

Observation e545fb63-b739-4379-885b-1e155e47775a · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.412866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.292479Z digest=sha256:5ccd9dd49149cbd3c0ab63efa36137b391e3827f9f99ae006e63077ab68f55aa

Observation 162c306d-abff-46dd-a029-c6edbdfa953e · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.391181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.297381Z digest=sha256:cd375b04a725d991fb62f9259def0647fbaed983cf3394d35fb12a3b94a0baa4

Observation 16c684df-b327-4f82-8bb9-7e17c22b4528 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.302280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.302280Z digest=sha256:1f4a69051cb3b201fe09f02ac9f1f1ef1fca0e2764d1f6151d94cb7f5a9e8753

Observation 2828e8b0-ff6c-4883-b663-885ba785378b · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.371372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.307062Z digest=sha256:92b6c7432964def2aee56131f69aacee87866b0073e9b54c2d35b37200b5f4f7

Observation 90ff68e9-ab1c-4fff-90a3-784606db2304 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.311995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.311995Z digest=sha256:afef514cec85c9318d2b826216c9eb2d0368bc56afe84716cc685a411b9fc0e1

Observation bd4ad7ce-5871-44c9-9b2f-0bc008ee470e · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model VideoChat: Chat-Centric Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.316833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.316833Z digest=sha256:7816779ba8feaf0ce76cb4c3078649730b45880dc592ed2f920b3563e548e151

Observation aea73c1b-d24c-42f7-9447-41c286b980ea · outbound

This paper cites Video-ChatGPT: Towards detailed video understanding via large vision and language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Video-ChatGPT: Towards detailed video understanding via large vision and language models

Reference 23

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T00:33:21.348872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.321745Z digest=sha256:29e5e361196968435d29409e4c7b5c05e7a7084cfca96128e2017a7bb437218a

Observation fa43ef72-8084-4d98-b812-6fccee3b5343 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.326679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.326679Z digest=sha256:9221fa9aacb4e59054b254961add0eaa40526404be0fd9848adac63b647bada7

Observation d03987c5-bb12-41b9-896b-94b8880e4c21 · outbound

This paper cites Video-LLaV A: Learning united visual representation by alignment before projection.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Video-LLaV A: Learning united visual representation by alignment before projection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.331593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.331593Z digest=sha256:b61347a219de94fec214d14932117a2b33fe71aa4920b0fc64a1f5a66644c537

Observation e40c7e43-5d2a-4814-a945-19f891f788fb · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.341323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.341323Z digest=sha256:d329ad139fc59d3dafb78a7293732ff9b3bce87931032e0d8be71a8f009fb54b

Observation 09ba4a5e-31cf-4a00-89cd-cc489715c8a5 · outbound

This paper cites SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.346344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.346344Z digest=sha256:2e4d82502f4a280048d6e41e4746ffdd2e470d754a320056ae11e143f813a0e3

Observation a75a3f1e-d172-4cc1-a04d-c8d055653d95 · outbound

This paper cites Slam-omni: Timbre-controllable voice interaction system with single-stage training,.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Slam-omni: Timbre-controllable voice interaction system with single-stage training,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.304020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.351418Z digest=sha256:5aa8370f6121abdcfd6d3723dc5bcba7b3b9549e819ced6994c232a8b69ace07

Observation 7e87629b-a43e-499e-b0e3-1baac53bd90e · outbound

This paper cites Robust speech recognition via large-scale weak supervision, 2022.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Robust speech recognition via large-scale weak supervision, 2022

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.278885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.361775Z digest=sha256:ebd3b8b1bb90325a5dbbc864e06ad02b67481b5455c7216d8b690631e6e0e560

Observation 20ba2d0d-0fac-499b-a766-14073c675d5b · outbound

This paper cites SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.356530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.356530Z digest=sha256:dda458dc8c161dd142a69442bfe3e5f7bc23d7fbeb96ba338ffcf98ba7ddbec4

Observation 76a7d486-85e6-46e4-a154-e4a7113a5e5d · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3451–3460, 2021.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Hubert: Self-supervised speech representation learning by masked prediction of hidden units.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3451–3460, 2021

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.371514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.371514Z digest=sha256:d260b42d8c7d6e9ded6fb7fc88cd8cd48a96b10fd0374624e4b82d0b396aa08b

Observation 364e34da-2e2e-4130-ae6e-9ecf434d071f · outbound

This paper cites SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.366716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.366716Z digest=sha256:da4f5f2b3614e1a3b5aed1b5cf3da9d5dfe4e2b1030e0217ee76f24733abe00f

Observation 40a80281-7f94-4d0c-a57c-851ee32bd298 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.381385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.381385Z digest=sha256:006fd34108495072c6a12274e8dbe2a280b32c141595ce9cf5203b2b38ae8acd

Observation 1393f3cc-df50-425f-a3c8-4714f50a3c8c · outbound

This paper cites Speechtokenizer: Uni- fied speech tokenizer for speech language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Speechtokenizer: Uni- fied speech tokenizer for speech language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.255218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.376928Z digest=sha256:83168ab99e09909bab0e1916cb9c8d44ed4cdb43262ab2db8a228ea7290e5a6d

Observation 49e0837c-b280-4bda-9abb-ed83e6a41216 · outbound

This paper cites an unresolved cited work.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.390752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.390752Z digest=sha256:c429873f94dbc7c7d2278b0fee9497e30c0b0e9ecfdc36807c9ac1779dd90dc1

Observation 9b572f75-01bb-4bdd-8a94-1835bf8faf8e · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.212324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.386288Z digest=sha256:458e2f2964f6a016c9dba19f9917ebcef1a9b20d829a428165f96b0a4b00b16a

Observation 6907a3a9-4e83-4aa4-a288-2d6d4560034d · outbound

This paper cites M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.400418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.400418Z digest=sha256:7c91537f58924ac7f3e67fc747b0a7dc35aa5ea0496e0e47f4e7264e9c63083f

Observation 79c360b9-12cc-47f9-8e9d-f0c67c6d7099 · outbound

This paper cites Megrez-omni technical report, 2025.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Megrez-omni technical report, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.191840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.395633Z digest=sha256:4d87677cc2edba39703982e45601717b56c5b06b7ad63f4f67edb2c89a857e31

Observation 3f5707ab-84f4-4a61-aa79-9f94835d864c · outbound

This paper cites EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.410942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.410942Z digest=sha256:8a2a9b460d7eefed877d2b7085fb6997140fc1083d1e9ce125ae5e9fb9440053

Observation 49d2bc70-bf19-47eb-9c8f-432f1887f8fb · outbound

This paper cites Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:33:20.191604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.405533Z digest=sha256:cf99dc51b46378826e2aaf0ffc73ee7f90d02e30a7e0713c19743e6613c64d39

Observation 5f8eeb56-c015-44ca-b3ee-76677e515931 · outbound

This paper cites Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.420806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.420806Z digest=sha256:318a032c30d28465749c441bd773924966a65bba3f2ccee19d484bb90c852c6e

Observation 2337fda3-cc8c-4241-b296-8134f263461a · outbound

This paper cites Openomni: Advancing open-source omnimodal large language models with progressive multimodal alignment and real-time self-aware emotional speech synthesis, 2025.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Openomni: Advancing open-source omnimodal large language models with progressive multimodal alignment and real-time self-aware emotional speech synthesis, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.416428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.416428Z digest=sha256:810397cbbfa73e55dc48a50312ba22de32d9234ebbdd02483dd5b196ca778b05

Observation 9eef4a63-d46b-45e7-b69a-d26e6226ab37 · outbound

This paper cites Stream- Speech: Simultaneous speech-to-speech translation with multi-task learning.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Stream- Speech: Simultaneous speech-to-speech translation with multi-task learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.430154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.430154Z digest=sha256:b705c2a6f46f24603818c7868f928bb0ce2352072b2a487683778f20bdae19d5

Observation aec48e8a-01c8-456d-98c8-0c62539af473 · outbound

This paper cites LLaV A-mini: Efficient image and video large multimodal models with one vision token.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaV A-mini: Efficient image and video large multimodal models with one vision token

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.172880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.425639Z digest=sha256:20842f612d13156a51cf51bb634824983986a3710f8b46d28bb49affadd2b80e

Observation 7b3b7fc3-c782-4dd2-9b12-c42d5f2977ea · outbound

This paper cites Future-guided incremental transformer for si- multaneous translation.Proceedings of the AAAI Conference on Artificial Intelligence, 35 (16):14428–14436, May 2021.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Future-guided incremental transformer for si- multaneous translation.Proceedings of the AAAI Conference on Artificial Intelligence, 35 (16):14428–14436, May 2021

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.153594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.439367Z digest=sha256:5a88b2955c83272c02f82cd920dac637dc451937402ed3f322662479be23366e

Observation 0d4c36c6-dd8c-4162-a85d-e910384fbe28 · outbound

This paper cites STACL: Simultaneous translation with implicit anticipation and controllable latency using prefix-to-prefix framework.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model STACL: Simultaneous translation with implicit anticipation and controllable latency using prefix-to-prefix framework

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.434764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.434764Z digest=sha256:e8ad65c74348beded56a5542f8a6a8e9bcd78146a3bfee692acd7bb6a4275da5

Observation b886f9d2-8782-415d-98a7-ca81cdb87e1f · outbound

This paper cites Information-transport-based policy for simultaneous trans- lation.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Information-transport-based policy for simultaneous trans- lation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.448868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.448868Z digest=sha256:a76e8edff23cb562e6aeaaf0ae7ddb6b5ee49bc3e975961979f17368158610a6

Observation ed288ccc-6968-47fe-9e67-c6d5fdd2ceaa · outbound

This paper cites Universal simultaneous machine translation with mixture-of- experts wait-k policy.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Universal simultaneous machine translation with mixture-of- experts wait-k policy

Reference 48

Resolution
verified exact
doi, observed 2026-08-07T00:33:19.693942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.443755Z digest=sha256:1a0c1657c484f6d13440dd5c532d7fc471a1872c0d32d41e8583b05504c5b16e

Observation 030d498b-c61a-4151-8daf-91be67b987b6 · outbound

This paper cites Librispeech: An asr corpus based on public domain audio books.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Librispeech: An asr corpus based on public domain audio books

Reference 49

Resolution
malformed identifier
no resolver link, observed 2026-08-07T00:33:19.458294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.458294Z digest=sha256:68d5b374f308e9c274db77f23dddb65f8a629c2e47c916e1a61ccb493b4c1b9a

Observation cad4eaf6-cc97-4b89-922c-a8f33a02d995 · outbound

This paper cites End-to-end simultaneous speech translation with differentiable segmentation.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model End-to-end simultaneous speech translation with differentiable segmentation

Reference 50

Resolution
verified exact
doi, observed 2026-08-07T00:33:19.662287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.453507Z digest=sha256:152565266d140fd00c3eb8d570f1e9bfe5d3ffc4a8602ccc8beca100e0a88f15

Observation 9ddbca23-aa03-4346-b39c-71a5b8142f41 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.133409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.467939Z digest=sha256:3e55bc55aa4f942222ce102698c305bb8a366672ef31f6598a067939df09366e

Observation 8b54d6d9-9aaf-4136-a244-4645ab435966 · outbound

This paper cites WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.462794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.462794Z digest=sha256:bf1aea60f659bf913d55a378af258681ddf3689855f1ff32e2715887cc6cc3ee

Observation 47529764-308a-4bec-98c9-0cf8a7c46da4 · outbound

This paper cites Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.088297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.477513Z digest=sha256:1e71293104493a998671af52c83b149dc48c4be7c45e773ac2c1ff5ea17ba014

Observation 15dbb48e-f8d2-4abf-a228-c45ec74ff014 · outbound

This paper cites Hudson and Christopher D.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Hudson and Christopher D

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.109643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.472638Z digest=sha256:b7cca5c0279363bc34931be9545eafc418f422a7f9d10fdd0f306366ac917ef6

Observation b4715a3b-9719-4bfc-8ba2-92d0d69d98ad · outbound

This paper cites Towards vqa models that can read.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Towards vqa models that can read

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.044743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.486601Z digest=sha256:de363e51894a8a80804c9bc2e685d3bb956a773099f22aca2724b864e7af0ca6

Observation ce3a6731-d55c-4499-acf2-7f72fe21f344 · outbound

This paper cites Learn to explain: Multimodal rea- soning via thought chains for science question answering.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Learn to explain: Multimodal rea- soning via thought chains for science question answering

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.065463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.482086Z digest=sha256:709f39d4dea5778037a57ebf60fba148f0826d49055d1a8d05dcfa38a83a8cba

Observation 9ab882af-001f-4b7d-9ed5-96fe76b00cae · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.496435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.496435Z digest=sha256:495f8f5cc6d4bb8af500ab1ebbcf0bed6cbee6c266cca56a8bb5902c3d6455b1

Observation f8475a16-8b23-45cb-a4c0-978f7933fbb5 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Evaluating object hallucination in large vision-language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.023280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.491754Z digest=sha256:0b810c245513c23c788c8e37d87623359f825c4610a0633bbccf8b123617e1d9

Observation 6f30e2b9-f600-4094-b497-c145b6453040 · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Seed-bench: Benchmarking multimodal large language models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:21.004498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.506706Z digest=sha256:47e20a06f4ab66e6c5b5606eb82b8c1b91328ffd9f96e99973e2488c210ea617

Observation a170ddd9-daa1-4f37-8cb4-3e6b391b00ef · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model MMBench: Is Your Multi-modal Model an All-around Player?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.501199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.501199Z digest=sha256:55c77e7d87f058ebb88d451323a3f3c69f5a58b4b844f7fc6730f96efd29df35

Observation 0e0579f0-f8fe-4228-ad87-283cd2471329 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities,.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Mm-vet: Evaluating large multimodal models for integrated capabilities,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:20.972598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.516384Z digest=sha256:4ff77af42b5336d8679d13902b9fc45730b97b6b687be8753bbdb205b6fbe523

Observation 783463fc-6429-43de-aa46-6cdf7ab10579 · outbound

This paper cites Visual instruction tuning.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Visual instruction tuning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.511589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.511589Z digest=sha256:05cc48525480049b7aed8ae8463d7f6cc4e92406f94d92a284181b9e5a411e8c

Observation 5a224011-d38c-4c0b-8e36-661891981140 · outbound

This paper cites Semantic parsing on Freebase from question-answer pairs.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Semantic parsing on Freebase from question-answer pairs

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:20.935198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.531289Z digest=sha256:6fd5fd8622110d85e9d7237aa9395f4ad93d649cdb6b84eac5e595ca2b5411df

Observation 3af33f61-c3d4-48a9-bffb-067b3f7454fe · outbound

This paper cites Visit-bench: A dynamic benchmark for evaluating instruction-following vision-and-language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Visit-bench: A dynamic benchmark for evaluating instruction-following vision-and-language models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:20.898083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.540625Z digest=sha256:05ff55c0be72f6e4d233bd15be79032e528bb10468cd7f63a6d502163808c9d4

Observation f5487bc2-ab3e-4bc9-b9d6-0d435ab04a66 · outbound

This paper cites Spoken question answering and speech continuation using spectrogram-powered LLM.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Spoken question answering and speech continuation using spectrogram-powered LLM

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:20.953856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.526568Z digest=sha256:6d2b21008544aa1597933874b58ca2cf648aea9d4d4ddd3e7790539ed9215339

Observation cf89375b-0d49-4cba-8248-1ed91245a399 · outbound

This paper cites Improved baselines with visual instruction tuning.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Improved baselines with visual instruction tuning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.549650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.549650Z digest=sha256:154968040c932730f7fea4b6c14f3d64504efcd45b616802d61a60c7c9db0645

Observation 40133eb9-41ff-4f21-979f-69e2e2a4f8ce · outbound

This paper cites Introducing idefics: An open reproduction of state-of-the-art visual language model, 2023.URL https://huggingface.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Introducing idefics: An open reproduction of state-of-the-art visual language model, 2023.URL https://huggingface

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:20.854697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.554791Z digest=sha256:0f6d9eb07449e1a2009b473754c1d21b3468a8a29e3bc629787410a04b5bd0cc

Observation f2e0cdb4-aef5-43c3-a9d0-eae9f60b6d8e · outbound

This paper cites Textually pretrained speech language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Textually pretrained speech language models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:20.835643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.559359Z digest=sha256:85a951d604c53bf1250cd126622e0fb48ef8e8df9df0f937ccea771010b21b46

Observation abe7db92-287e-4465-9595-8a98817694c7 · outbound

This paper cites BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.545187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.545187Z digest=sha256:8b6e58ef8a69ff2c0396021ab468dc33caa0e40455da6cdfea4b78cfe0d4cd31

Observation 1cfecd7a-cb94-4f44-ba50-8e016548ba32 · outbound

This paper cites The Llama 3 Herd of Models.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model The Llama 3 Herd of Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.568881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.568881Z digest=sha256:2fe4815b599b47db8e359879b5a82953b0aa913fa87502339767e728134f976d

Observation a119a8b9-f13b-4860-8699-792aabae8383 · outbound

This paper cites Sigmoid loss for language image pre-training.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Sigmoid loss for language image pre-training

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.573384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.573384Z digest=sha256:030135f0a24a37366a8e41491914c0354c38e2349f386bd1a7fba232b9c5ea90

Observation f5e67540-2a28-4247-9b5c-61fabd5ceea5 · outbound

This paper cites LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.577773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.577773Z digest=sha256:d48f633f5a62f691b2eff0ae7c987f909448ec806f76feba6b820e772a455cef

Observation 387faf8f-1424-435a-b860-dc09ef65783b · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.563774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.563774Z digest=sha256:267ad64adbfea8ef01f72ebb05f398da5ab01648cb92519a0b6c1b877eaf7ef5

Observation b913fec9-ae1a-4c18-bfd8-c3f2b8284f66 · outbound

This paper cites Enhancing chat language models by scaling high-quality instructional conversations.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Enhancing chat language models by scaling high-quality instructional conversations

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.586908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.586908Z digest=sha256:d68776c47895388763f020e7f665cf335a61dbcc40760e13f8bf3fbda59f4ddb

Observation 22b2724e-132d-4889-9947-a15400715054 · outbound

This paper cites does not allow traveling to the second floor.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model does not allow traveling to the second floor

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.591854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.591854Z digest=sha256:6139f8a00955347ae89c0e1e18900464e92a46ff73959dbb70478b7706180c1d

Observation 7d2c1a34-e4c4-48c0-b0cf-f4ee13d27e4e · outbound

This paper cites Scaling Speech-Text Pre-training with Synthetic Interleaved Data.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Scaling Speech-Text Pre-training with Synthetic Interleaved Data

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.582482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.582482Z digest=sha256:64db3afe49603af08ffb26365935f669e186fba5cb72b839d02391230a055de0

Observation 8c062e03-716e-412a-8997-ad76e25128a1 · outbound

This paper cites URL https://aclanthology.org/ D13-1160/.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model URL https://aclanthology.org/ D13-1160/

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:33:20.916950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:33:19.535826Z digest=sha256:602d248cf905df2fc47c82262195fe96d2207aa738c72652904857dd4dcbd122

Observation 8ede076c-5894-4de9-8ed9-890ce438eb60 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.521264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.521264Z digest=sha256:5bfdd3d0c1da12e759882e8638e239d5bd120f028a5980b5b44ba8510dbb69b3

Observation cba64abf-8252-40a7-9f0b-0e0ebac5daca · outbound

This paper cites doi: 10.18653/v1/2024.emnlp-main.342.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model doi: 10.18653/v1/2024.emnlp-main.342

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.336391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.336391Z digest=sha256:1b884cefcdfbd40e761e728e05386684aae91d263c042dc969385d5e20adfed7

Pith citing papers

Observation 286ba852-3dfa-43a3-9a01-4abb28a5dfab · inbound

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models cites this paper.

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:13:52.168111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T02:12:55.170296Z digest=sha256:12a21aa8c352ab98c058411ae8c8f31a8e79fb57ccab6d6b44c1d143274332ee

Observation 20d11b5b-424b-468b-98b5-45b0698792ac · inbound

PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory cites this paper.

PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:59.686225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:05:40.242925Z digest=sha256:76457ec881c0ef75436d99a536b32540be28db88ecc7dc947784febd63a12404

Observation b3a8cf48-e086-43d3-bc8e-60ddc7414b0b · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.634543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:5230695a092381ec19fb418ecd0dcf2bfa87c57bc459ba9ae609237ab52e84d2

Observation 95af9d12-ec95-40ab-8db1-5aad1f844823 · inbound

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning cites this paper.

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:38:39.603859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-03T16:37:06.384435Z digest=sha256:720dc06edb95819209081271a8a4d047a769a11c88447f00b2587ebab002c178

Observation f28a4588-daff-4cac-8c3e-8020f4f57764 · inbound

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory cites this paper.

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-11T06:35:35.951554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T06:35:35.951554Z digest=sha256:0e30c19e45bcdf4607bb36584337644c6af7b8259dda54f399da96ca68d52a04