Pith. sign in

Paper Citation Record · LEDGER

LLaMA-Omni: Seamless Speech Interaction with Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 62 inbound Pith citation observations for arXiv:2409.06666.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.06666 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 62 of 62 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:47:39.105733Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:29:41.322265Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8ddcf9c2-c033-4b24-87c2-3fbb7dc8e4a2 · inbound

VoiceBench: Benchmarking LLM-Based Voice Assistants cites this paper.

VoiceBench: Benchmarking LLM-Based Voice Assistants LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:50:13.971194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T00:50:13.841689Z digest=sha256:cbd541cc11ad448506c03c4eb264dd852c137de02bc04c203d07f84daca5042a

Observation f6d2f7bb-0212-41b2-821f-8761c3a01feb · inbound

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot cites this paper.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.496972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:e77845a19efee9abf98224266e338addb38bedd41af961148ae87c1181f5c54a

Observation 31b3f180-c57e-4f57-a86e-2b8d1f4efb8f · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.661305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:ef188192c9fc8b0804dde45b87704d4e595e989c56020694f554d312b7be34a0

Observation 7dc3286d-0748-4cc8-888e-601fe4d49c74 · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.105733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.105733Z digest=sha256:d3b84ffe914b8d56fba23c1cd47c4c1a8a8979df5d0d4a3d2ede8896bb050cf9

Observation b5002ec4-157b-4ce9-983b-9e285c585c7e · inbound

SparQLe: Speech Queries to Text Translation Through LLMs cites this paper.

SparQLe: Speech Queries to Text Translation Through LLMs LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T22:05:23.804113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:05:23.804113Z digest=sha256:f0c3daec966c58edb43586d8658d9f7cd5617e40a137722ab3b68280aa53ece3

Observation d54a85f0-f0cb-403c-84be-bf3409720cca · inbound

A Preliminary Exploration with GPT-4o Voice Mode cites this paper.

A Preliminary Exploration with GPT-4o Voice Mode LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.322585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.322585Z digest=sha256:e930be0bf98da5fe9d8702f8497efd8bf64b6e77ad93c096031899165a43758b

Observation 2d3ae2fc-a0b0-4740-8024-10f7dbc5bf8e · inbound

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction cites this paper.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.342720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:05360e91f4ee5aee40ee0b937918d8497e46efb3cb8cc36426dd72427303a096

Observation 033a91c0-6928-4fc0-91f1-3a1bdd811f0a · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.234568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:479c4b2fce86911c44a73cb723f6d3bc187f0069f0f19dc352790f323f86b331

Observation c07ba69b-84a4-4031-a186-ef95e4cb43fe · inbound

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach cites this paper.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:33.966621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:33.966621Z digest=sha256:ef9862b6b3f4770a7783a76d0d2fc67767578af7dd8d4e7d559b6469a2bf5b06

Observation 507fb53f-1c52-4fb5-81b3-7edc9f0ec53b · inbound

S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models cites this paper.

S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:36.938679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:36.938679Z digest=sha256:82ff6baf243a2c1672a81cce187cb0b2b0cb0c783403ed466cdafa9adaf01969

Observation 180b0dc2-a501-411a-aaf2-b27ae8112244 · inbound

ModRWKV: Transformer Multimodality in Linear Time cites this paper.

ModRWKV: Transformer Multimodality in Linear Time LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:34.644120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:34.644120Z digest=sha256:84ba5d6f3ead14e6e87dd36c415f8e948efedb5aaf671150034b73d62d8a0083

Observation dc9e1a5e-b5fa-4f82-8711-d91ee80fa622 · inbound

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models cites this paper.

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:01.147173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:01.147173Z digest=sha256:e14f0cf46d11af4516229bce7ba8c1b6dfa9cb301ca79b643664a2ef1ed758e4

Observation 0d13692c-21c8-42a9-bc12-cfe3f818b5c3 · inbound

Speechless: Speech Instruction Training Without Speech for Low Resource Languages cites this paper.

Speechless: Speech Instruction Training Without Speech for Low Resource Languages LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:16.069890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:16.069890Z digest=sha256:beb9cb7624c6c4b845c7ae3adcb7c5749c1273310ae60da8e4aa55d5e1903ec0

Observation 9ca4f5b7-4a33-40d1-aa74-c22a6e2656c5 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.644063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.644063Z digest=sha256:52e2472c99d59a6866fca4539c46e9d8bf7c071fc9a8ad1f72129780edd9812a

Observation a6c1b79e-1f1d-42bd-9d60-646fde200bf9 · inbound

MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding cites this paper.

MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:33.562854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:33.562854Z digest=sha256:01ed15d8b8a54e300fb2ccd725543796fccb35b83e5f5b22fc675e766a48c1ab

Observation 132089e7-bcf9-4ff1-ada2-857a733acffe · inbound

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling cites this paper.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.598942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.598942Z digest=sha256:9cbb761e9f238034d534cb83784e34801100678760f96e0f7eb9258aabe1bbbf

Observation d802cbdc-6ea6-4981-8b07-cc34f1544fc9 · inbound

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction cites this paper.

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:30.307140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:01:30.307140Z digest=sha256:4a0c3c88855092294b31b57ba976b0aca3b0dad21f2ddfec6a4b015b3cd72469

Observation d3b73640-6507-408d-a905-1fce1cd35252 · inbound

Universal Visuo-Tactile Video Understanding for Embodied Interaction cites this paper.

Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.344367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.344367Z digest=sha256:2ed77f596e5f9b6ea3d212635626dfe709a052d3e45770e1fe7b03d7bc270f58

Observation ce5592c6-903f-44ff-ae2d-861a45d3c6c1 · inbound

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems cites this paper.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:31.966356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:31.966356Z digest=sha256:99625d338d5f21e37784eabd7d2ba2c8f31bac9dc661ee480169978cc48e5188

Observation 9277ae40-93a8-402b-a2fb-6f5c118d4c39 · inbound

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction cites this paper.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:35.040733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:35.040733Z digest=sha256:fd249f86170e9d8cf09920b4a6aaaf065e9990922737d95ca531586922dc26c8

Observation a38ced2f-e6bf-428d-9ab4-bc652ad434f4 · inbound

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion cites this paper.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:39.960930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:39.960930Z digest=sha256:0161cb728f340fad8281e1c5e36ed9f14e44606a94354e1cecfd80e619c24896

Observation 3a38c437-6768-4ac8-b6b1-defd92aa8332 · inbound

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant cites this paper.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.008363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.008363Z digest=sha256:8f44c552de1539268e7fcd57c8a0a31b8f0c841235b30c7dd086437cd39e1a4d

Observation 076e0b54-13c8-48fa-aa92-1a3b5b45342a · inbound

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment cites this paper.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.924179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.924179Z digest=sha256:c8f3d02497ca65181a90cb0bf7ff4a6cdac3701e350792cc9b2c34497210703e

Observation 1985f09c-790f-448f-ba7c-612b0d17d029 · inbound

Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models cites this paper.

Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.921978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:08:40.921978Z digest=sha256:c450c3b743933098b29d1743b27f7ad3e87d3eefb6861e9d6e704c5a57ee18e7

Observation d734e001-cd5d-4eee-9cd4-638a837d3b22 · inbound

KERAG_R: Knowledge-Enhanced Retrieval-Augmented Generation for Recommendation cites this paper.

KERAG_R: Knowledge-Enhanced Retrieval-Augmented Generation for Recommendation LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:23:47.250006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:23:47.250006Z digest=sha256:ce2553cbd61ad1dd9e6056089ddc9b727b2d94d90d640a5966f4231fa1618959

Observation 60989bf5-f0e4-4c5f-ae7d-85c95a1d3c30 · inbound

Unlocking Speech Instruction Data Potential with Query Rewriting cites this paper.

Unlocking Speech Instruction Data Potential with Query Rewriting LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:21:32.233590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:21:32.233590Z digest=sha256:74da8dd7f8b5e4e002feaa83f00328a20580ef8a88bb8dfae793933444a6bb47

Observation c396d604-debf-4f00-846d-96cf0fba6658 · inbound

AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation cites this paper.

AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:47:55.871612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:47:55.871612Z digest=sha256:9960e0fbc5163fd87d5440fe0602ccf8d0f51c09cce85f7d164bd3a195dc5f35

Observation f9ecc50f-fee8-4c79-bd65-941214304f01 · inbound

Personalized Socially Assistive Robots With End-to-End Speech-Language Models For Well-Being Support cites this paper.

Personalized Socially Assistive Robots With End-to-End Speech-Language Models For Well-Being Support LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:09:04.402999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:09:04.402999Z digest=sha256:1037976fb7298f634eea011d737d1e333e5fc5a1a4ca618db44f8fca018a073f

Observation cdb8ac00-ddad-47a7-abd9-6780ddf68f6e · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:51.023035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:779f453e7fec296c7c50de43a70d903b46d5d3867c2f1ecb0504302dec004c9e

Observation 275754a4-9a8c-48fc-8896-d67dea81c1d9 · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.291383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.291383Z digest=sha256:1100c2b79da2f058b3a293b77484f342c104b39e0fa2ecb570f405fce94942ba

Observation 65f20b7f-5213-4063-be88-bbe4978d02b0 · inbound

SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding cites this paper.

SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:56.043822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:56.043822Z digest=sha256:9972c4bd5a87c0f8198ff04a8361da7814af208b82993e7e1c170efec382edec

Observation 60e45a49-bb4c-47c9-8812-24061e498963 · inbound

Dual Information Speech Language Models for Emotional Conversations cites this paper.

Dual Information Speech Language Models for Emotional Conversations LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T21:45:23.907075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:45:23.907075Z digest=sha256:769837001fc02ec95cd6860cf07131068b52c8323eaefe02d524fbdf40338f69

Observation ee928dbd-a44f-4c45-9c7c-8826b58e7890 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:12:54.167660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:4587c4caf195dad559ac89fff2ae9fc8bbb2e7353d7604c8b2fb71691a98fb6d

Observation ddc7832d-be84-4b6a-a021-89c7de4e3344 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:05:30.612129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:71ed9451f9a7b3678c638c6452e19a025b13b30fc9f6860ba72fe17e339d6848

Observation ff509164-128d-453b-85bd-ae5fef3a68fa · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:51.856505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:51.856505Z digest=sha256:2f34d523dfe9aea4d7b9bbbeaaa7445ec749aa29e2e756df96d9d92729cca073

Observation ef60b9df-dd01-4cfb-a286-ad7713ca6c18 · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.513736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:87f3f23ecb3dbb7688213a520e307b6db81801000757ca94c85379eed5a72d1f

Observation 91a0b4c1-b79f-45db-8bae-2e63819c3e44 · inbound

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations cites this paper.

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:04.355944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:33:04.355944Z digest=sha256:6196d9a2839c3d294e1fba65b946a967ec5eebaa76ebd4cc51a1f89d83d0bf2d

Observation bca84238-c1e2-4017-a014-0ee3f5b7d255 · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:01:24.361876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:88a26bd4866ae80c1cd56bc703eeb163508ac31ef02393860cd3c104e0409daa

Observation 55ce28e0-4813-4353-8197-81b50a8f88de · inbound

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues cites this paper.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:44.086529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:44.086529Z digest=sha256:2c782034e649d9f4bde762bc9e830a93dfb1c5344baf903ca4652e2f00b16875

Observation 73d20c5c-f39d-4b87-a946-54358ef0a2a3 · inbound

Same Words, Different Judgments: How Preferences Vary Across Modalities cites this paper.

Same Words, Different Judgments: How Preferences Vary Across Modalities LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:36:33.038978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T19:31:51.996480Z digest=sha256:f2f9edb52be9956699197df6cd2365c4acf395a836c22b07a0a772eb8730df98

Observation 941f052b-a146-45b8-b2c5-7596fdacdd7f · inbound

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision cites this paper.

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T18:42:05.886184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:42:05.886184Z digest=sha256:047cc1ad31991696b723f3e6d77d4c2d41c61f582226d84fda1a0f77e4dbab79

Observation 8cf905bd-13d5-4852-88e2-5551808ab626 · inbound

Controllable Accent Normalization via Discrete Diffusion cites this paper.

Controllable Accent Normalization via Discrete Diffusion LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T21:25:19.253843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:25:19.253843Z digest=sha256:27ecb2a1ce2e7acdc3a040ea630923f38488439c52aa189d31801317b9469c8a

Observation 18bc937b-afd2-45c0-be4a-79329bb0cbe7 · inbound

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model cites this paper.

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:08.488015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T15:34:50.848124Z digest=sha256:000837ba798cf38d7614299c60564b98e87ca87745088c9605087bdffa9187d0

Observation e57d5363-c6af-4b86-8681-b2d59a1be6c0 · inbound

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization cites this paper.

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:45:44.079790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T18:08:16.051117Z digest=sha256:f02028b3f7d9e6ae6d21a27100b090ee2eddac92fce6a02cb40977c50be5a5bf

Observation be8c5c01-a85f-4883-87da-1846317b008e · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:46:15.262510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T11:00:52.196039Z digest=sha256:e46245abc987f7419809d63d34963809716a3e4882e66b085431fd393149a901

Observation 7cb90703-4bc7-4fcc-9550-4a20e3920691 · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:49.611468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T00:49:26.507281Z digest=sha256:727d231ef5ff3c012fb33ca9fc6d61a3f087b5244c805c456f8fe33652867719

Observation cf1b243e-3b0c-46ae-8eb7-1a1e3436288e · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:55.866632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:79d83fde0d5bba8a7eb9556f65650bc301e5a8e668797feb3db30c313cb4466f

Observation df2398ca-422d-4bb7-b0a5-6b4af5991141 · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.969176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:f786384c0e2fe8d52e2e7d9cfae1016ed293931c68b880a9d2637e3631fa92c8

Observation db74770c-bf98-4781-9492-e175e48af0a3 · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:43:55.027484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T02:41:13.583493Z digest=sha256:053145da1abf0858c457aa451aa8186f31cda6d78ec250b2c7f1f17a0e02913a

Observation 4838ffaa-416f-4089-8ce4-244f17c814de · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.361577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:9178fef824d5b863610d462209529fa7bd3456bff9457e519111a7140bb46b10

Observation a6d8c1b2-019d-468f-ad1a-814d5b2a6a22 · inbound

A Survey of Audio Reasoning in Multimodal Foundation Models cites this paper.

A Survey of Audio Reasoning in Multimodal Foundation Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:09:24.368254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T02:08:06.976461Z digest=sha256:b58ab36cdbd21b27205cbf12a63e16b53c6c893bba6a105c95f3c7f23732061a

Observation a43b7391-d3e0-4f63-9793-c00a05313c7a · inbound

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens cites this paper.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:25:59.872963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:f548c8211cbd582dd5027ed0bcbe2eb3df2614b90e814682bc9095c6acd6573b

Observation 65b49c1b-dac7-4476-b120-6205597d2b6d · inbound

Audio Interaction Model cites this paper.

Audio Interaction Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T10:46:52.413130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T04:57:05.062465Z digest=sha256:b6cd4816ad845aa81e97a69e2f1384f7081a3fcf2ede2e065abe88828f593aed

Observation eb3d44dd-5b67-4797-96c8-e90f5dc44120 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.662681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:bf1e1f520e96ffeb7a0d7ea6fc71247e0d941cc0a0858c1b86dc49641cb6d6b1

Observation 084cd552-38f0-4ee4-a5de-dac922514b7d · inbound

Steering Where to Listen: Instruction-Based Activation Steering Redirects Temporal Attention in Large Audio-Language Models cites this paper.

Steering Where to Listen: Instruction-Based Activation Steering Redirects Temporal Attention in Large Audio-Language Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:57:44.819997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T11:29:36.012349Z digest=sha256:b0bf66a2baf87c1b5874f52d534ecc8a4282ac3521e7fc103c56847092feeffd

Observation a53e1ce9-d567-4f85-9ed0-d0ab6702f9c2 · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:12.825126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:01f5c12041a8609cff7287b793d439573b0312ddf090c0c91da0f6e74cf0d591

Observation 6f7c5947-006b-4b80-9cfb-cf4596cd38eb · inbound

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead cites this paper.

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:29:41.323872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T11:42:44.035902Z digest=sha256:077a8c69ed8b3b398a12aecdcac3d50f2b3f8d54e41ca8ee1f70914a8063d46a

Observation 2ae65a48-4f96-4a77-8557-222645a0aa62 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 215

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:55:42.281653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:622666a679bb31172727ef3d185070e6fc08e378e8f2ca0d8aa9f301ca849815

Observation 6199ec6f-8f95-4bcb-a4d8-b464e14d29db · inbound

Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving cites this paper.

Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T08:09:44.957032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T08:09:44.957032Z digest=sha256:f328e08f2d38bb0c42e349cfe6dbc2a6cd7eaf5999eeab274c9570383f83c4eb

Observation a130e201-e309-4fdf-bb17-135aa45b6047 · inbound

TokAN: Accent Normalization Using Self-Supervised Speech Tokens cites this paper.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:2b9e60f6f8de0a09daebd747360a5f3e30b11d3ec3a787d6fa3b1728b42557a8

Observation 788183f2-e31e-4971-a137-16cb22e0c543 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:13.127289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:13.127289Z digest=sha256:83944c4f97ffc6c4d36c0aa1978b85b8e59d59c5723789e327b0d50d54d696d7

Observation 5e6763b9-2544-45e1-82ae-e2b4c8cbe348 · inbound

Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO cites this paper.

Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T01:56:11.344932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:56:11.344932Z digest=sha256:b3a96db3cc87c69eb6886d46762153bf8092bec2854303f39b31765eb9128a6a