Pith. sign in

Paper Citation Record · LEDGER

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction

As of 18 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2606.09186.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.09186 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T15:22:02.107863Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T16:37:24.214630Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact8
  • verified fuzzy0
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch16

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7de6918a-8e30-45ad-bd8b-ba362a98e847 · outbound

This paper cites FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.977710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:f1a22ecc64bb63ed931714d72ff8e285c19f4ce9d1fb3e6e2c1c9eaf97773d99

Observation cd1e3c9b-e5eb-41aa-9191-8d12d5f9c043 · outbound

This paper cites Qwen2-Audio Technical Report.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Qwen2-Audio Technical Report

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.967180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:49398fd14f1a7853c021b11ed300cb79ebac557a637a7684368beee984701ae3

Observation 6c6575e6-57dd-432a-bd55-12a5ff1fc009 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Moshi: a speech-text foundation model for real-time dialogue

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.974880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:0ab7858a8ee4d53bd33431e71156fef87f90d87ab8b1e1160905149afee4032c

Observation 616e258a-1439-478f-b0b9-83aa4e3fe9a0 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.990361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:49512ca94dff1058a59ec62014987b37574a07096f04f5e92b4e092418473a87

Observation 98fb4e5e-9dad-49dd-88fa-f8daedc1774a · outbound

This paper cites an unresolved cited work.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T15:22:02.107863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:e49f74cc499d726077f35314f74240134e80b3b4083bd22c462abc1bce74103e

Observation 47e44b60-ac6b-4a55-9448-b864553eae8e · outbound

This paper cites Baichuan-omni-1.5 technical report.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Baichuan-omni-1.5 technical report

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.948374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:3b0cffec51d7ac4a12bdffc1ff0d9d88eecdb3850fc002a3fb8bbb84fcb67ff5

Observation f3ebd363-802d-46ed-b0c1-22820ed014bf · outbound

This paper cites Towards Better Instruction Following Language Models for Chinese: Investigating the Impact of Training Data and Evaluation.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Towards Better Instruction Following Language Models for Chinese: Investigating the Impact of Training Data and Evaluation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.996815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:e07555d9bbdc5af2980e258202dafd39f63becefa1cf6ca67215a16074248c50

Observation 520d023e-46ba-4a4d-9dc3-fb488f1c84ce · outbound

This paper cites Kimi-Audio Technical Report.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Kimi-Audio Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:27:34.934034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:05e6e50de038fa9b3f89cb3e7cac8058f85c6a56a77df3b257a6647a00db8e38

Observation a3aada79-a49b-4af6-b4b1-98a1a95fe510 · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.964427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:e36c68a754388b0c1df89d99f50f6598232a960fcba8b6dd8f0081af37db9d56

Observation 95185f7e-3b6e-4b9c-8acf-470044852c07 · outbound

This paper cites InIEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025, Honolulu, HI, USA, December 6-10, 2025, pages 1–8.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction InIEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025, Honolulu, HI, USA, December 6-10, 2025, pages 1–8

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T15:22:02.107863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:bd548960a24efaade932b628ec21d0cbe5cc0054449ddac413731c11839c1c90

Observation da5bd0c2-fcbc-43ca-b841-79eaff1e1b73 · outbound

This paper cites CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.994077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:941c85734ff6e2443db1dcbb8a400deaa5e7481115eaff4aa3d40810a64af5c1

Observation 719b2b1d-0a84-4af9-8df2-24f3c64f0019 · outbound

This paper cites an unresolved cited work.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T15:22:02.107863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:9d68730183298f4edfa674ea8d9aa985be371d7945a039631b3929d50ae2cbee

Observation 1b0a8841-f964-46db-991b-6d8ad15baee3 · outbound

This paper cites Generative Spoken Dialogue Language Modeling.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Generative Spoken Dialogue Language Modeling

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.967134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:abdac0ff106200f05f7bc6d7eac587ae4907076f3798ccdbb24cab26ce238dcf

Observation 1424cab8-256d-4987-901d-2e21e602c5b8 · outbound

This paper cites GPT-4o System Card.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction GPT-4o System Card

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.925449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:8882028207471b739a41cf26e6de20b902bd91c20c03167382f1c7bd65d79632

Observation f486d2a1-263a-4b1a-b8bd-ad83515765f5 · outbound

This paper cites In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2015, South Brisbane, Queensland, Australia, April 19-24, 2015, pages 5206–5210.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2015, South Brisbane, Queensland, Australia, April 19-24, 2015, pages 5206–5210

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T15:22:02.107863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:f1808953ebca5701a677e05a461c05f791ba4892605ab06fa19e902aebb544e9

Observation 1540ff3a-9419-411e-b956-0665f08be33c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.951605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:78d27326f9987c2e435f4109fbc09e38fb19a0483971e0e9755ae0d7ca652ad5

Observation 4c428786-99d2-4a8a-a407-84414cd2cada · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.984043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:dd1c3a881e82f84ab56f7613c1545fa8e66c62c5d8a8ceb056ecc439d45bacc0

Observation 45001c07-ce4d-46b9-982f-cc7b3b081eaa · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.933850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:8ae8a3efbb17d9286abcb1f249e07d0f447f988088567c2183ca23756397876d

Observation b1ba49d5-88d7-4bb1-a437-86645d3b9ac5 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Gemini: A Family of Highly Capable Multimodal Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:27:34.964777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:34e8f88c32cc96e19673ea60b5c285d8106d0c185efb09c213b63481812926fc

Observation 6075f2df-37ec-4ec1-9a62-71f21f480513 · outbound

This paper cites Fun-audio-chat technical report.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Fun-audio-chat technical report

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.982969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:b84e388ce21da6606cbc61ef04a17e62ec4fd75113109921fec5baece7309819

Observation 4fbddfca-2ba8-4d45-9932-5fc094535bec · outbound

This paper cites A Full-duplex Speech Dialogue Scheme Based On Large Language Models.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction A Full-duplex Speech Dialogue Scheme Based On Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.976859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:af7537d74daef169bc96c61712f872b6702c2355a42cfb4c2f2c7fe786552a58

Observation 677c4174-76b2-4d78-a095-13dd3c96d46b · outbound

This paper cites Mimo-audio: Audio language models are few-shot learners.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Mimo-audio: Audio language models are few-shot learners

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.987343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:ae990c903c83dda83350bf7b20279875a03d9c5c9d0b2007848d54cc4425b742

Observation 6a7e3871-3f95-4a75-81f1-ed871e2981fe · outbound

This paper cites Mini-omni-reasoner: Token-level thinking-in-speaking in large speech models.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Mini-omni-reasoner: Token-level thinking-in-speaking in large speech models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.980125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:be54a69f5c4a6044caa1a1749abda09040b3ae52ea377c79b0fa6438e1efbbc8

Observation 1c71c09e-45b4-4f30-ad3c-be4af4142bce · outbound

This paper cites Qwen2.5-Omni Technical Report.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Qwen2.5-Omni Technical Report

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.942903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:69e45962e3e8c22569866c6b6f25def85b57839f0ecae607f89a9ef6c80319ce

Observation 5a806667-89b4-40b5-ab16-1e6a32e8592b · outbound

This paper cites Duplexcascade: Full- duplex speech-to-speech dialogue with vad-free cascaded asr- llm-tts pipeline and micro-turn optimization,.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Duplexcascade: Full- duplex speech-to-speech dialogue with vad-free cascaded asr- llm-tts pipeline and micro-turn optimization,

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.980639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:37f833a119ade437de01f8fa8ecd7370bd7cb90cb329371f7c7c2c03fe464dae

Observation 830949d4-e30d-4e7b-b065-6f7a62fae384 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.973705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:a27970354e86c1c72eff8fef99ac703c6f37dcff998fcab1dcd1e90b21ac2fff

Observation 030fe35c-2674-446b-b5c0-b3fc2ffab60a · outbound

This paper cites SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.936721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:31d414530379bee708c26b1a96e9c690d7c7b29679cf518810c16da76fad88da

Observation e301429a-ae85-490a-b8ac-a531e9cb089e · outbound

This paper cites WildChat: 1M ChatGPT Interaction Logs in the Wild.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction WildChat: 1M ChatGPT Interaction Logs in the Wild

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:27:34.931397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:ce9b04a4180433d606afc56eb5482833919e6225621fd2d1cc79a4d6c75bbb69

Observation 569fe23b-d8b0-46b2-9cab-4d122910491d · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Mlvu: Benchmarking multi-task long video understanding

Reference 29

Resolution
malformed identifier
arxiv_id, observed 2026-07-03T03:27:34.970446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:c8c83b76e13213a46a0e4e67e5cd5025afb526d247eaef20cf62d0c5509d26bf

Pith citing papers

Observation 89b5d4f3-5ebc-40ac-b54e-9fcb33ce1ef0 · inbound

Voice Memory for Agentic Speech Recognition cites this paper.

Voice Memory for Agentic Speech Recognition DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.214630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.214630Z digest=sha256:90a35f47b9e35d365a9dfef953a050571b575dd4ba977258fcdeba2aa1ebcfa3