Pith. sign in

Paper Citation Record · LEDGER

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

As of 22 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 9 inbound Pith citation observations for arXiv:2506.23325.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23325 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:51:37.442979Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:37:36.541957Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T19:57:20.080166Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3a7c3175-3b79-49de-852c-4d8effdb4a49 · outbound

This paper cites GPT-4 Technical Report.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:33.740032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:33.740032Z digest=sha256:d7bdfe42015944793ec3e811e56d13dcb57475c10295ac6a845512754c6d60e6

Observation f8961e6d-6e17-4a20-8b2d-45e4adda53b7 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Moshi: a speech-text foundation model for real-time dialogue

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:34.365210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:34.365210Z digest=sha256:b47297f2650acc4ac33a9204f526f883c77a58b609f989286a4c6a159beb52e5

Observation 7985bc61-4f57-436e-a2a0-4ee9939435d8 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:51:40.699479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:51:34.446168Z digest=sha256:012eab3b083548280a2c8761b7562fd2a179451cb995a7ffc611185a430e7179

Observation 270fa8d5-7e1f-4bbb-9997-77ee954e258a · outbound

This paper cites Gaussian Error Linear Units (GELUs).

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Gaussian Error Linear Units (GELUs)

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:34.904183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:34.904183Z digest=sha256:7dcfc2418d0552b13a1624f21124c199e948840b508cdd1ae1abfbc71680b729

Observation 762248e0-df10-4ea3-ac8f-a499ded69049 · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:35.040172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:35.040172Z digest=sha256:f2716d704a524669c207d4bbaa4023229bce95ba650cdb5693a18e8dcdd40d41

Observation b0168cde-54d0-420f-92bd-c76a6392a6e2 · outbound

This paper cites Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:35.263754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:35.263754Z digest=sha256:2dbd3842c01c5c1025d10dc9a6735441bd6d7020d23a56e94d8629ea6a28fcc7

Observation a1156adc-ae8c-422e-89ea-20d90a542ded · outbound

This paper cites Decoupled Weight Decay Regularization.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Decoupled Weight Decay Regularization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:35.500706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:35.500706Z digest=sha256:397a658eb861fa44c9a3da45940659134475db311917c181560af0df4cf1f483

Observation 155da7b2-5768-40fa-a998-96514840291a · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Robust Speech Recognition via Large-Scale Weak Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:35.718054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:35.718054Z digest=sha256:595bac5f2e1279e226e5e9a5893200bb41cb818f3b2b82fcc171800327f1dd6e

Observation fb92abd5-8c07-4867-89f9-8af0cdb9ea6f · outbound

This paper cites A short-time objective intelligibility measure for time-frequency weighted noisy speech.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs A short-time objective intelligibility measure for time-frequency weighted noisy speech

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:51:39.721809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:51:35.979072Z digest=sha256:85f3c5f83a45bf1243170ec2bb14758522b0be2d99390c9bfa524f9f6a2c0418

Observation dc09ecda-a6e5-439a-862e-d73db211cc55 · outbound

This paper cites Unified Speech-Text Pre-training for Speech Translation and Recognition.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Unified Speech-Text Pre-training for Speech Translation and Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:36.079536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:36.079536Z digest=sha256:b1fd0eabd62d207803bf086053d28c3941328a213c962d133d706cd61718a530

Observation 294e4a4a-e8e3-4cbd-8a99-078047d1f6d8 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:36.176279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:36.176279Z digest=sha256:0ff9c90a4b47e27a0cb48e50b062068d7b68fe9c8f00f8cf00496dbb94027aed

Observation 429258a5-fdab-4c66-94c3-0078deb08521 · outbound

This paper cites Towards audio language modeling -- an overview.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Towards audio language modeling -- an overview

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:36.281078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:36.281078Z digest=sha256:d14fe8e9393f8e97036136f90d4f854f29fc87f0b853278b9d188130d2ffabd3

Observation 8b56d155-f599-4703-a5f6-958f59f252f0 · outbound

This paper cites BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:36.381830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:36.381830Z digest=sha256:22a0dd2a9174908f03c54b5dd85e67030b3ef9eea34c4aff91789f5b2a93a870

Observation 15e22bf4-f8da-44ae-a13f-7632dbec2ef2 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:36.537269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:36.537269Z digest=sha256:49578bbb945b869f08c1a6d47eac0269d2a98f03c26328eec1919178dd510fbf

Observation 7ed8cb77-4e20-4711-b181-e625ac32b5ec · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:36.668481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:36.668481Z digest=sha256:05e1a903ff7f7ef131da079b67d9151265ef0f386a1fcac0db2de201e4618555

Observation f5a9656d-2e5f-47eb-8e74-63c39c1aca01 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:36.827023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:36.827023Z digest=sha256:547ac75215214d1eb452d66febb57f4c3faf688f862c625e0c079503f0cc196e

Observation 371d1574-2abe-408c-bb78-31d12a2e3bb2 · outbound

This paper cites SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:36.963958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:36.963958Z digest=sha256:7b5bce22315fccd0f5192e28b5ef11dddcc299caa3fb4c30c13e6f813950c02f

Observation 7a83db08-20cf-467e-848d-1b4ae1d757d6 · outbound

This paper cites Each adapter consists of a 4-layer Transformer with a hidden dimension of 768, a feed-forward network (FFN) dimension of 3072, and 12 attention heads.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Each adapter consists of a 4-layer Transformer with a hidden dimension of 768, a feed-forward network (FFN) dimension of 3072, and 12 attention heads

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:51:39.368558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:51:37.038424Z digest=sha256:d146dd4bd90b2f966905a44ec8e45e131105961895729e2f582c90468a55d9b3

Observation 23b56d3a-3cde-4619-98e9-be2d5bfde1ba · outbound

This paper cites an unresolved cited work.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:51:39.100411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:51:37.085694Z digest=sha256:aeb72490b5e4b80614892250d63f8dd9d71bcb4f514c7739ac07d32fbd0228f2

Observation 3073077d-3b1a-420a-ad5f-7af336c8af43 · outbound

This paper cites an unresolved cited work.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:51:38.871784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:51:37.144477Z digest=sha256:202f0e21f2eaffc4d87ed37672b5ccc2a49d17c74f52efa9e47a49c2acf831b7

Observation 1daa3878-b8a5-4747-b686-b452da095a2d · outbound

This paper cites an unresolved cited work.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:51:38.588276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:51:37.218866Z digest=sha256:591f06351c3ece630c4a5456c4c77e0d1e3035de683933d9e63798f4118da047

Observation a0053048-6b25-49ee-a7e9-30d98e65bab3 · outbound

This paper cites foreign language.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs foreign language

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:51:38.436659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:51:37.333548Z digest=sha256:b076e48e7de783790da49a094a74448265f855e2cdc361588332fae25a053c5e

Observation 39b1a29b-a44c-43b1-9016-faed2ed8dca4 · outbound

This paper cites an unresolved cited work.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:51:38.167437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:51:37.442979Z digest=sha256:5154226689f9456ee9b7bbfa2684fe6a17a596ccb5f256deb35c44e191f9edb6

Observation 16507183-20b4-4104-8119-91ffd070184e · outbound

This paper cites Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:35.862909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:35.862909Z digest=sha256:63fe8615a32e39c21e90216848f388d70de092122801bb607b29c2f952dceebf

Observation 390e4e58-c607-483b-91c9-a4730a8080ce · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 2006

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:51:40.385701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:51:34.773420Z digest=sha256:7ad260ccc6cc2c8e8f415d0ff748160b36d0e294f6d5ac71dffba2537b060b69

Observation b842017e-59b4-4f45-a7c8-c50e5b39e8aa · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Librispeech: an asr corpus based on public domain audio books

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:35.623910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:35.623910Z digest=sha256:776a3fc45f916c5bb11c57d8b7862c2162c45693ad56ccdb354223077e940470

Observation d2c491a1-b8c1-4810-8c63-8db7a26621c1 · outbound

This paper cites Kimi-Audio Technical Report.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Kimi-Audio Technical Report

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:34.544816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:34.544816Z digest=sha256:c7dcfd478b383a53916e92651e93568bd70a2f2e262b72964f446d23e5501562

Observation 000e211a-5b01-4ec4-a8c4-6173650183aa · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:51:40.055159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:51:35.819686Z digest=sha256:dc929366eadbe2d1bfee64a99db4471f7037d522540f465d5d780780fb625f8f

Observation 7959a35b-01db-44b9-aa0f-728901b69685 · outbound

This paper cites High Fidelity Neural Audio Compression.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs High Fidelity Neural Audio Compression

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:34.229293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:34.229293Z digest=sha256:6c72f721a8a72ea7d22ad92c0154f60b6763ab8fc07917dcb74fc7a2dc7e3190

Observation 19ff56c0-2839-46fa-9e34-4f9a7d54da0f · outbound

This paper cites doi: 10.1109/jstsp.2022.3188113.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs doi: 10.1109/jstsp.2022.3188113

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:33.872073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:33.872073Z digest=sha256:8acf6d07095d5538b08af9b0ea8a39bbcc6bc8fb7284c78238b53298e002f00d

Observation b81e8bff-6915-42ea-a251-eb248a06c311 · outbound

This paper cites Sparks of Large Audio Models: A Survey and Outlook.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Sparks of Large Audio Models: A Survey and Outlook

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:35.148862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:35.148862Z digest=sha256:81da7ae72c910957f69f00bcd1bdbb5068221cd059342b1bb5784d989b7be32b

Observation 891805ed-fab5-43a8-8948-f4277db70652 · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:51:40.999128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:51:34.082036Z digest=sha256:55b2facfecf58d1721e2b0d7b1a5876b63a98b0997c312b4cc5be83cda59b57c

Observation 5b8dfb06-9aca-4ad5-ab64-86ced8db648f · outbound

This paper cites Flow Matching for Generative Modeling.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Flow Matching for Generative Modeling

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:35.387371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:35.387371Z digest=sha256:32bd5ca72fd58512ed502382801fc5969a28495fbd51dd8960abcbf2b5eb997a

Pith citing papers

Observation 827d65ca-cae9-4b43-9f3c-912e7370c341 · inbound

AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation cites this paper.

AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:59.738880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:59.738880Z digest=sha256:641a060ba38fcb4214a1d9e8095d958e62df20ab7fff60b15b21098aa5084be1

Observation 68ea7121-7178-46a9-853f-f6eea513ebac · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:24:56.099066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:084b1e9dc467a23f03dbb4dcbd18003e676e83fa2d1776741e6fc45c035f5568

Observation 8868fc17-160d-4d5d-af6a-863c6956a272 · inbound

VERITAS: A Multi-Agent Co-Scientist for Verifiable Image-Derived Hypothesis Testing cites this paper.

VERITAS: A Multi-Agent Co-Scientist for Verifiable Image-Derived Hypothesis Testing XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T21:34:30.078813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:34:30.078813Z digest=sha256:f46dfb4cfbbb716109d2d9ca565a3fa2ef91beb028a2b89bcbebcb0d93f6542d

Observation 022ec4b3-f695-4420-9a06-905f81ca4397 · inbound

Why Your Tokenizer Fails in Information Fusion: A Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization cites this paper.

Why Your Tokenizer Fails in Information Fusion: A Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:01.898524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T14:49:51.481873Z digest=sha256:5269708e83452a64c15021e245e1e6ce7fc04a8c03b4ec91b9566ff3a3012a61

Observation 919d90bb-c069-4a92-9139-adee8d952b25 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.020011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:b88b0287555d144a341734bcd19059a5b1e73a28b284f341988c336effc1e0d2

Observation 594065c4-4506-4376-87c6-54f77b9c3ff2 · inbound

Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation cites this paper.

Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:23.233947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T01:11:50.733585Z digest=sha256:3c9123ebe023c343830ed9c83ee54254b1811cef69aeb83acde8cae8c71d3dad

Observation da31b60e-9d12-4c5a-ab33-74f6270c59ac · inbound

dots.tts Technical Report cites this paper.

dots.tts Technical Report XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.081577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:424914d6e3391302ee667b4006ab9ce4269766daed79a4510aab60a709e7e3e3

Observation 4c74f3f1-d424-4b63-a725-7f1931abe1aa · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:55:42.290110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:61121ac8563bd077e880f497038445e3a8f96df4a15d18b67b1ba9f1b0d7bd4b

Observation ecb525e6-c774-4269-aad9-03ca55301624 · inbound

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances cites this paper.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.541957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.541957Z digest=sha256:48f60ff925d91f6564b94a1416399fb93d4df89d0207d621a3f5c020639d75dc