Pith. sign in

Paper Citation Record · LEDGER

Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2411.00774.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.00774 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:41.303521Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T02:49:25.043605Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation abae069a-dbed-4db4-87c3-2b4187cdda19 · inbound

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot cites this paper.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.520398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:546b8de6383e15e6a6c213abd17a0075032734f76921071d18b95ac084d2c306

Observation b647c8de-ec50-4927-921c-acc2855f6f16 · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.680358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:197b88d160b3a525eb35107a2e7b6b808825fb11c5c8173d9613461dace790cc

Observation eef0cf8f-ad35-490b-9f52-6c2b723ccec9 · inbound

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction cites this paper.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.390382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:baa82215d2c5870d6052db0ec2fa18a75dd0175b2835930ebc352adba888ee90

Observation a3436f6e-4265-4c1a-9431-8b3e1c4da8e8 · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.083009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:422f0ae054e066bf3639df0da59c57d73f8e80bab6316bc4e55d7d207a58f0bc

Observation 70d29ff1-c4a9-4617-9f1a-5ef617c2d357 · inbound

Beyond Words: Multimodal LLM Knows When to Speak cites this paper.

Beyond Words: Multimodal LLM Knows When to Speak Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:31:36.003748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T13:31:26.532921Z digest=sha256:c517a715bc71c1729caafc836f03bdf11cc4fb2d5d8c077592529717190f08cf

Observation 33b497fc-fbcc-4410-a8e5-9be4634c3e0a · inbound

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model cites this paper.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.303521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.303521Z digest=sha256:67bae6f5ce369a9e25a4c2954d05926fd0b9989579c3713b8a55ff86851652e5

Observation fd4d79c4-b7f4-466a-bb6d-a80d9ec59e96 · inbound

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant cites this paper.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.253202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.253202Z digest=sha256:7dfd2d612fe6778353aa7b4b016e3d33195517a227e7b18c5b5c4de02c82129e

Observation 0e806ba7-755f-4f05-b107-ee54be2e91b8 · inbound

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment cites this paper.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.015033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.015033Z digest=sha256:496b6d675a17cb52a185bb9f40b4775721010de043089432167aa15fb69362a2

Observation e40c7e43-5d2a-4814-a945-19f891f788fb · inbound

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model cites this paper.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.341323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.341323Z digest=sha256:c81aef29f6300c8d51fefe10d236eff6c5a5779d9318481a4328ee3ff3c164fa

Observation d0e3b8bf-90c4-49c6-8219-cfe9ef22e744 · inbound

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction cites this paper.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.469864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.469864Z digest=sha256:052c8918743a55711e2b04eb4adf4a4debcaf40ca6300c11b5eb6643eb522939

Observation 532a21d1-7a9f-4c26-a582-7550813991d3 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:51.063110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:0b4e268076aab54a785528446d909b6093116bb81695eda667999daf391f9b3e

Observation 9411af84-8233-4e5b-af70-fcd682a869a0 · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.306448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.306448Z digest=sha256:fbe1b17f0adc439d3daf450d698230677bf1350549703b8b30972b3e4236bc65

Observation 9506fd3f-2eeb-458c-b029-ed36eae169eb · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:12:54.079300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:ac2b9809e7aa3d4ed5359e765b6553873a6e65dd3847f4602f125f0cedd61a4e

Observation cbe65bb6-1dab-4fbc-83c0-aa5d6f3ad7fa · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:05:30.811981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:bf482ce922652332069c21b877412cc59b7480076621300d27bc07dbd268104c

Observation f0ab0b96-75d1-446c-944d-1a3dc177c3d3 · inbound

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations cites this paper.

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:04.388843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:33:04.388843Z digest=sha256:05f203b901509917ed167651e6047381aef1dfbefe85442b362f3d2cd31fd40e

Observation 050ac82d-14d1-4658-bac8-7b4904c2fa6e · inbound

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages cites this paper.

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:31:37.205947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T16:27:37.596817Z digest=sha256:14777f41407c6752fd9e8817fa89fdf11c05111418d78ddeaadc5bef33905d9f

Observation 6e2ce822-5c73-4964-9132-5967e2ff9bbc · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:01:24.286498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:db4cc192181cb7504952cffe387068c4e56c6fa131f1a1a9adc085ffa10630fc

Observation 875ecdf8-0960-418d-b24a-1bc98129c738 · inbound

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues cites this paper.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:44.360493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:44.360493Z digest=sha256:a365978dabebff2d1c2dee16d412c0b11bac8490934bce468d8039215a557296

Observation b6c7d276-875f-4385-b79c-ff404d291883 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T08:35:18.308768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T08:34:56.898815Z digest=sha256:e04cb65b640562508bde96b82cf469bf6071f7f2e49511181fcb15d7f2948b04

Observation 3c672184-6bf8-4797-8a9d-457a07d83018 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T10:44:07.690636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T10:43:27.176535Z digest=sha256:7b95c45b9375239b7f59c124a4c800eec357b8ca9d9ae0f8c4e6d196958928b0

Observation baa5aeb6-f915-4bf1-9d7f-579f1722c107 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T22:57:11.059962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:57:11.059962Z digest=sha256:06691830fd0d1b637042b6bb0ad694a70b1e51fcf7cccc75c0dfa8a59308030a

Observation 406b4d52-04ec-4799-83ab-23ac10707f45 · inbound

Sharp spectral estimates for free boundary problems arising in plasma physics cites this paper.

Sharp spectral estimates for free boundary problems arising in plasma physics Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T19:56:33.793167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:56:33.793167Z digest=sha256:3fe7ac479e357b7dece4d073c0c30891f595412338efd93e15257caa40daac0a

Observation fa7b331c-ebc2-4724-8893-422a48f620ac · inbound

FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection cites this paper.

FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:15.977586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:55:09.141906Z digest=sha256:2b678386cb444ded46ed7886bb3edffe3d6c397c20d63f877060825ae4971b03

Observation 64dfc8f1-7a60-4ab7-99d9-8346edfbfcaa · inbound

Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge cites this paper.

Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:56:06.773428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T13:19:44.156822Z digest=sha256:bd2cebe2d9f1f65af475da995f301beb742d92c942c9c8812f0de3d69c521003

Observation 64976695-458e-4845-a071-2d77b9ec5bc0 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:55.782660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:6d1f0215a67450f91956702274decfd3243e4d1aad063726f45bdb6356cd78e2

Observation 50d7ac62-7679-4c69-bb0b-3563ebc80c57 · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 123

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.633916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:00b90ba6e69e42578e8f5e228c82f3e4083d251c1c7a02ea176391c693f23fc6

Observation 5f1e38ba-87d2-4a20-86df-7c8f1a7fa4cb · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:43:55.070044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T02:41:13.583493Z digest=sha256:d995d6d4c54de0fc7dcfdd89253a21c1747ea5e0d8242bd0a3fee7978e91144f

Observation 9ba951a5-3065-44ce-8f9f-83ddbaf3159d · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.380968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:b80e805b4c0dc14f6cc9c857699eb511a00603113606dc35ecf6c65588195256

Observation c7955dae-ab99-44bc-9b38-f2357ccca974 · inbound

Raon-Speech Technical Report cites this paper.

Raon-Speech Technical Report Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T08:22:44.347941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:22:44.347941Z digest=sha256:bfc8667c7c1baf0b6ec1e8c9711e364b7b3cdbb70b89b8531dbd574aca1d9c32

Observation 82ceda9f-63db-438f-8c2c-4edb2cd1c888 · inbound

Learning When to Think While Listening in Large Audio-Language Models cites this paper.

Learning When to Think While Listening in Large Audio-Language Models Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:43:50.760010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T18:37:34.409802Z digest=sha256:b32e3260a77afc365432d2add3cb807c546dcb5d2793083b3c9b188d8d1cadd2

Observation 7903bc9d-253b-4979-aff2-1d48af51bfaa · inbound

Audio Interaction Model cites this paper.

Audio Interaction Model Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-02T10:46:52.400863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T04:57:05.062465Z digest=sha256:c7dca69937f925cb85e4193f72a5d47254abd78eec0882ae50e5804ebe614e98

Observation 886a5fbe-a2c0-4e68-a2b6-86fd427837c2 · inbound

IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems cites this paper.

IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:37:06.807408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T23:42:38.203116Z digest=sha256:9fd3c2c6f7bdb3399dd7e4e71d930f10cbbd48733f944207d931d59bbc4d6619

Observation e3854dc8-bc48-4565-95da-21b99a907f74 · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:12.835252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:273458e724c3858655b472b608affba693031479c18228d89e35502af506b38e

Observation 606dd4cc-93fe-49d5-a83a-511298c5da8f · inbound

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents cites this paper.

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:28:39.533604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T05:32:27.371597Z digest=sha256:d605dedf6fbdb5d6d57593a16cca703c02e84d4fff1db02ca609a46a0e35b6bb

Observation 1321db71-8e75-4f38-a674-e6fb7383c948 · inbound

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents cites this paper.

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:43:51.429171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T05:01:44.202712Z digest=sha256:ebe6c92d70c80c3c6eb4b379a644fc1fbb19770986e5bcf781a84d763903e6d8

Observation 64a7b10f-e728-42bc-8c56-6c2e0ddecfdc · inbound

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine cites this paper.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:25.045148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:8bc2cacbd6031e4001c5b1fa670505280e5e7fa5587df51524f9e546c31dce48

Observation 24b875d5-c0d1-43c8-8c33-da00620d2ab7 · inbound

FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars cites this paper.

FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:44:19.416518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:38:28.718164Z digest=sha256:0a827b1bc95a26c3e6cb0f4aa44134ff7b21ee14a369660310d116639a793ca7

Observation 3cbc1c14-2319-4db7-a69a-b13171bb03a9 · inbound

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation cites this paper.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.847968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:d26e465b24e1f9b35341f2b0b342c90970dcc610daec7c90d5f170f08e7883aa

Observation c5891d88-aa53-4c4e-a6e1-1cd840b0eb77 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 216

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:55:42.220803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:fa48cd1642e21a03b14d31fd33a9a257134c7be4c4053265e859296429180016

Observation 986e7776-14a4-468f-bf4e-a2445abaa5ec · inbound

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning cites this paper.

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.611230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T14:35:17.004680Z digest=sha256:e4cfb8128529383981fb4e2fa44b70e8b63774773fbea5d25161f98935f8ca44

Observation b000e732-f9cf-4a4f-9c51-3b275345b487 · inbound

Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions? cites this paper.

Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions? Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T13:57:31.306452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:57:31.306452Z digest=sha256:6cd6180ba7a0dc0249c4831fbe124eb1fb9ab56d050455dbdc04e45dda4cbbf6

Observation b26e331d-2228-47ad-a1d3-b11dd57b5500 · inbound

M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models cites this paper.

M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T13:16:11.540208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:16:11.540208Z digest=sha256:6f39d006c464731ef2f06413a48bed84bcff385b138725a24050a012af91415a

Observation 52398a3e-4123-4ff3-866e-7a4699707149 · inbound

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents cites this paper.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T00:30:10.821180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:30:10.821180Z digest=sha256:563d2fc59b33e510c41ff4173be5ac4488aa82f1e9b6050bf1ea62f23bc6410f