Pith. sign in

Paper Citation Record · LEDGER

Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2502.17239.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.17239 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:38:37.251024Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7ac78b78-52a7-4d26-af96-90e2f944f2cf · inbound

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies cites this paper.

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:52:16.736897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T23:51:43.934329Z digest=sha256:a84802afcba8e438e11ef7522978e759313d8bea9c4a9276728c2ac87802567d

Observation 93c544ca-6fdf-4cf3-9a36-01f9219931fa · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.223308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:8379a0009e9d9b8139b7789b669e566ce6c2aa2e5232ee6f7c8e48f9e8f94429

Observation 3fe741c6-d071-41b1-9000-f47f5998c79a · inbound

S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models cites this paper.

S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:37.251024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:37.251024Z digest=sha256:04a9ea159519b2cc0da49a615de0c95479e493f13a3cdc5b158e723a00042673

Observation a861450a-14e7-4ae6-b631-4ebdec81cd59 · inbound

Breaking the Barriers of Text-Hungry and Audio-Deficient AI cites this paper.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.822450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.822450Z digest=sha256:f554c42d45a486bc3a5ee14be1506abb2bc4cc23fedacb57fd76e3d8db42ed87

Observation be26c935-a7ac-46cc-8640-094e2ff6b65b · inbound

AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks cites this paper.

AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:53.227261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:55:53.227261Z digest=sha256:296012a73aa9c107e1452ff192a02bd41f45ba15a7d8db0b0649be7bf40ece48

Observation b0168cde-54d0-420f-92bd-c76a6392a6e2 · inbound

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs cites this paper.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:35.263754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:35.263754Z digest=sha256:90fe41f67977246ab306ad8a9d6f8b7906e44d421458b3a361fae229beededdc

Observation c5ec8a30-318c-426e-93a7-7007b19df62e · inbound

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models cites this paper.

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:42:44.716460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T03:42:44.523919Z digest=sha256:42e7402d1bd3bd6239e2ac4b8596eeb8224201fd42d6f6b6e6971af778414cc8

Observation 40b1610e-2ed8-40ec-b53c-be989790466a · inbound

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation cites this paper.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.020393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.020393Z digest=sha256:df2ff135ef15d2480653be614891678778e5a768abbde72b99341df69fc98c62

Observation 6eb69524-e602-42c6-9335-744ef968d22f · inbound

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents cites this paper.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.427752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.427752Z digest=sha256:9a1e92c1c628d319e18d1abc6b7cde91f60ce192d0ce635762da8b21177e3e17

Observation 6926b130-d033-4f16-badd-5e1c9729d192 · inbound

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages cites this paper.

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:31:37.190198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T16:27:37.596817Z digest=sha256:de0347f0aa678a8c3d87a8488fb1fce2248ed6881772177a61a4d470cb4fd11c

Observation 89d9be4e-f9a5-4f7b-a26b-0bc092ce3677 · inbound

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents cites this paper.

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T10:13:42.410505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:13:42.410505Z digest=sha256:2a376817c2dbaea7d9c32656b43982672cf4607075dad39eab159f9f764980a0

Observation 11e27bd1-6153-4f61-8b74-5684833c4c51 · inbound

MultiAPI Spoof: A Multi-API Dataset and Local-Attention Network for Speech Anti-spoofing Detection cites this paper.

MultiAPI Spoof: A Multi-API Dataset and Local-Attention Network for Speech Anti-spoofing Detection Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T18:00:28.526023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:00:28.526023Z digest=sha256:31eab5e95cc94f761d82b6d0da2341450dac991d43be218bae7dbb471041b0a0

Observation 94e8db46-1223-4201-b70b-90ca2fa41cf8 · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:39.457626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:39.457626Z digest=sha256:1601dcfc82155b476a0b2bd206b1c99c4e6d073a5cb8d0e844e8a705b527cb6f

Observation d6683298-010e-43a2-9606-620f3e4996e6 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:35:18.399022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T08:34:56.898815Z digest=sha256:1ad18848540076b8d8f9ea499d9ca3638b6692257b830512416b2a6faf4ff7f2

Observation 1619f4df-1231-4214-b6e6-ac818e748e8a · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:44:07.754424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T10:43:27.176535Z digest=sha256:23efebbb6eb95c747a40eca6213b5c150c31f518b8e022e8c26ad7429bd0c8e2

Observation ecce569e-f6b6-4ab5-a71e-af81032600c7 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T22:57:11.059962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:57:11.059962Z digest=sha256:1f3f9f6bd35f324ef2e5053e812c7fdf1f0d1f714d9a9f44364637f641b96f12

Observation ae31ae8d-13e0-4ce0-a579-01c2dac98406 · inbound

TiCo: Time-Controllable Spoken Dialogue Model cites this paper.

TiCo: Time-Controllable Spoken Dialogue Model Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:39:35.811984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T00:38:52.182973Z digest=sha256:e165c73ffdf0239bc05b49229ad5f8a915b0759139dcc2a9b522c6dac0f8d54a

Observation 7baf8c86-9613-4ee9-9ec7-cfa1175ab18d · inbound

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model cites this paper.

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:08.056194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T15:34:50.848124Z digest=sha256:5ec034b7c9a996fc60deed770f48ad9cec3bb62ff47833638bbc24e966ffb73d

Observation 61026f86-cd55-4732-8fe4-d45b601d3ce4 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 104

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.122055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:bd373af1f2b40bef63e99f9620451d82f12e8ebea51d961134f6fc4cbd96c91b

Observation 7ec77828-2a15-4d3a-83d3-590e8adb6a76 · inbound

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue cites this paper.

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:11:19.093344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:08:55.753359Z digest=sha256:1df30454847c07bd79809ae60af019c2b8c4c43c77637738260aa5c57c575809

Observation 6b2082f1-4432-4c22-9f48-8dafc6f09f5d · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.790217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:0c0c3903ac571d67b447814a2ba5a4ade706e0093a905e7e196ef3e2e7913c61

Observation bbbc490e-cd64-4314-b29c-33519ad62833 · inbound

Raon-Speech Technical Report cites this paper.

Raon-Speech Technical Report Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T08:22:44.347941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:22:44.347941Z digest=sha256:40c1fc5ea35c4f73a2db0fbfbbef908f670d7f1b1fd39dd00d4c777a32bb37b2

Observation 99d7e8bb-82e4-4a41-85f5-5c04d7ac008a · inbound

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects cites this paper.

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:52:26.925475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:44:07.669223Z digest=sha256:604a5790b4d69b6b2511e243d1069588a519fd0f6d8c21cd3ee12a61476533d7

Observation 181163d2-d0d8-453f-8e9a-015a68694417 · inbound

MOSS-Audio Technical Report cites this paper.

MOSS-Audio Technical Report Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:25.114734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:a604e0de2ff453cc8af5e40b4548cf6f368f274423cc8feea9161c9d37b87456

Observation 4fc400e4-c217-41ba-8839-4e5e8e252bb1 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 218

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:55:42.134456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:378def9096332429b7650da0bddf29d3e58d46ad79ff37f03fb4e1d108951bf2

Observation f7baa26c-ec12-481b-9480-c87ea1f3954c · inbound

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models cites this paper.

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T11:20:17.818345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:20:17.818345Z digest=sha256:2c65276eb06cf27fc11f363c186580cd3ecfec824dfc482b0f909736a43158fc

Observation 5e51c39d-defd-4a26-86bf-4cf89f94e98d · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:11.242181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:11.242181Z digest=sha256:880a67abc964deef9bdd39ffdcf99e387a1fa2a5f00e04ac15eafd27e3802562