Pith. sign in

Paper Citation Record · LEDGER

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction

As of 18 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2606.09186.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.09186 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T15:22:02.107863Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T16:37:24.214630Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact8
  • verified fuzzy0
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch16

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7de6918a-8e30-45ad-bd8b-ba362a98e847 · outbound

This paper cites FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.977710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:9d4ee6d1c894fcaa4f1eb34e0937e2c89e82c53d5029fd006c21c7588319bfb9

Observation cd1e3c9b-e5eb-41aa-9191-8d12d5f9c043 · outbound

This paper cites Qwen2-Audio Technical Report.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Qwen2-Audio Technical Report

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.967180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:e5f0672a0f14b79677d3e101b4866136911836fc1d9ff5bdc1f33f4dbe956571

Observation 6c6575e6-57dd-432a-bd55-12a5ff1fc009 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Moshi: a speech-text foundation model for real-time dialogue

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.974880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:6ba932813034d310bf37d99dcaa041ea51775651f4e65ec56fc9392f9afd4067

Observation 616e258a-1439-478f-b0b9-83aa4e3fe9a0 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.990361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:b5ee2af77a8bb481b35b68ae4260754cba8c07bca91be492a02527412925bd5e

Observation 98fb4e5e-9dad-49dd-88fa-f8daedc1774a · outbound

This paper cites an unresolved cited work.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T15:22:02.107863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:e49f74cc499d726077f35314f74240134e80b3b4083bd22c462abc1bce74103e

Observation 47e44b60-ac6b-4a55-9448-b864553eae8e · outbound

This paper cites Baichuan-omni-1.5 technical report.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Baichuan-omni-1.5 technical report

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.948374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:0a8ed54e22d488dc2f82b16e5ee2f636b8093325a149710081732ad943f35e4e

Observation f3ebd363-802d-46ed-b0c1-22820ed014bf · outbound

This paper cites Towards Better Instruction Following Language Models for Chinese: Investigating the Impact of Training Data and Evaluation.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Towards Better Instruction Following Language Models for Chinese: Investigating the Impact of Training Data and Evaluation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.996815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:b7e69a00f2c63d89c0bcab9ed2d06c7ee57baeac6532ea171213da0a0b11f6bc

Observation 520d023e-46ba-4a4d-9dc3-fb488f1c84ce · outbound

This paper cites Kimi-Audio Technical Report.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Kimi-Audio Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:27:34.934034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:866577009a963520ee213c718bce0766d0e0402ad9fe57350d9eb8b70336185d

Observation a3aada79-a49b-4af6-b4b1-98a1a95fe510 · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.964427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:b935ef75197f4a76d0e755533a39659f149aa438b97d1c101178c9b730a38b54

Observation 95185f7e-3b6e-4b9c-8acf-470044852c07 · outbound

This paper cites InIEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025, Honolulu, HI, USA, December 6-10, 2025, pages 1–8.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction InIEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025, Honolulu, HI, USA, December 6-10, 2025, pages 1–8

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T15:22:02.107863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:bd548960a24efaade932b628ec21d0cbe5cc0054449ddac413731c11839c1c90

Observation da5bd0c2-fcbc-43ca-b841-79eaff1e1b73 · outbound

This paper cites CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.994077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:776ead80d95248ae5b97103f8ea04a26db78697e792acb05b949a7646118ce8a

Observation 719b2b1d-0a84-4af9-8df2-24f3c64f0019 · outbound

This paper cites an unresolved cited work.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T15:22:02.107863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:9d68730183298f4edfa674ea8d9aa985be371d7945a039631b3929d50ae2cbee

Observation 1b0a8841-f964-46db-991b-6d8ad15baee3 · outbound

This paper cites Generative Spoken Dialogue Language Modeling.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Generative Spoken Dialogue Language Modeling

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.967134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:59adb768fecacc2cdbc66f36dff5f63d0ff0ad8e971c8509cdc972d1b98722f7

Observation 1424cab8-256d-4987-901d-2e21e602c5b8 · outbound

This paper cites GPT-4o System Card.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction GPT-4o System Card

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.925449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:85517e27d3f97e31dd0bd0bbfc049ed7f434f1e42e76a7d34c083624ef977b9b

Observation f486d2a1-263a-4b1a-b8bd-ad83515765f5 · outbound

This paper cites In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2015, South Brisbane, Queensland, Australia, April 19-24, 2015, pages 5206–5210.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2015, South Brisbane, Queensland, Australia, April 19-24, 2015, pages 5206–5210

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T15:22:02.107863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:f1808953ebca5701a677e05a461c05f791ba4892605ab06fa19e902aebb544e9

Observation 1540ff3a-9419-411e-b956-0665f08be33c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.951605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:8a5aa6e2cbdfc5b2ec3a4504923feea7f16d71032b53eeb3a535681438f94043

Observation 4c428786-99d2-4a8a-a407-84414cd2cada · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.984043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:c2697ea8024b4e733f32100c3845dab0faf44595f117dd8c67ab74fc0a6df47f

Observation 45001c07-ce4d-46b9-982f-cc7b3b081eaa · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.933850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:065a6b8be970931cb356070c7c052771a26297286d2704b79ed0ccdab575befd

Observation b1ba49d5-88d7-4bb1-a437-86645d3b9ac5 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Gemini: A Family of Highly Capable Multimodal Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:27:34.964777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:1323827680cf1f1ab668161a963a7afe9916737c038a009dee438f6dfbd516c8

Observation 6075f2df-37ec-4ec1-9a62-71f21f480513 · outbound

This paper cites Fun-audio-chat technical report.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Fun-audio-chat technical report

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.982969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:ab6652e7748847593b13911a903cc3076c3586364890470d9f403b152b331562

Observation 4fbddfca-2ba8-4d45-9932-5fc094535bec · outbound

This paper cites A Full-duplex Speech Dialogue Scheme Based On Large Language Models.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction A Full-duplex Speech Dialogue Scheme Based On Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.976859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:e2deca78b0565f115e0e10f2a60fe6c37d8fdd88eff11fe2e830315fa482af66

Observation 677c4174-76b2-4d78-a095-13dd3c96d46b · outbound

This paper cites Mimo-audio: Audio language models are few-shot learners.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Mimo-audio: Audio language models are few-shot learners

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.987343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:a49d93a26e97c7c70f67c358ce76c3ca9e2db09368c77e4d8c3ae65192f6c533

Observation 6a7e3871-3f95-4a75-81f1-ed871e2981fe · outbound

This paper cites Mini-omni-reasoner: Token-level thinking-in-speaking in large speech models.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Mini-omni-reasoner: Token-level thinking-in-speaking in large speech models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.980125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:5d017b05cbfd01759b26f84efed6f0884c14dba7762e036b1a53c5e7651e1eab

Observation 1c71c09e-45b4-4f30-ad3c-be4af4142bce · outbound

This paper cites Qwen2.5-Omni Technical Report.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Qwen2.5-Omni Technical Report

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.942903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T16:38:13.453678+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:260d7b0ec9584a34dcbb43f6bdda91d94b44bfda45eab8d6f43b493252c4c2ed

Observation 5a806667-89b4-40b5-ab16-1e6a32e8592b · outbound

This paper cites Duplexcascade: Full- duplex speech-to-speech dialogue with vad-free cascaded asr- llm-tts pipeline and micro-turn optimization,.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Duplexcascade: Full- duplex speech-to-speech dialogue with vad-free cascaded asr- llm-tts pipeline and micro-turn optimization,

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.980639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:800b327f453dd8d14e77f29baee406f77fce8156e0038165a05c848322b9c334

Observation 830949d4-e30d-4e7b-b065-6f7a62fae384 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.973705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:ca8df614627a505fd87c2c743c77df48cdee73105b97f36093bf45ea649f162c

Observation 030fe35c-2674-446b-b5c0-b3fc2ffab60a · outbound

This paper cites SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.936721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:61e50b38336be7f24b93945e770f38d58196b7c986d7e1b93382ee921bf8f4bd

Observation e301429a-ae85-490a-b8ac-a531e9cb089e · outbound

This paper cites WildChat: 1M ChatGPT Interaction Logs in the Wild.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction WildChat: 1M ChatGPT Interaction Logs in the Wild

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:27:34.931397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:5ee2d65c67ce31d72e6e64ae07a8c8d15d70f1c6e5a5295a9473469674153534

Observation 569fe23b-d8b0-46b2-9cab-4d122910491d · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Mlvu: Benchmarking multi-task long video understanding

Reference 29

Resolution
malformed identifier
arxiv_id, observed 2026-07-03T03:27:34.970446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:15c0e3bba14f5854732c6c91de1db2c6f93ff25029d26648dc5350c01f482e58

Pith citing papers

Observation 89b5d4f3-5ebc-40ac-b54e-9fcb33ce1ef0 · inbound

Voice Memory for Agentic Speech Recognition cites this paper.

Voice Memory for Agentic Speech Recognition DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.214630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.214630Z digest=sha256:90a35f47b9e35d365a9dfef953a050571b575dd4ba977258fcdeba2aa1ebcfa3