Pith. sign in

Paper Citation Record · LEDGER

Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2303.03926.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.03926 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:30:27.053709Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.393626Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c31f0988-d576-4e59-9964-a9dc3a1e4237 · inbound

The Rise and Potential of Large Language Model Based Agents: A Survey cites this paper.

The Rise and Potential of Large Language Model Based Agents: A Survey Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 272

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:47:52.719788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:47:44.152066Z digest=sha256:538a8a344c3e8aa893f0aa968ee371307e85a46ad74df56820c2dbf6c88ced28

Observation 979bebc8-0d3e-45fd-8ca7-22bb17eac15d · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:26:37.352136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:ccabeb62c7a6660e225210bee221d1efa7b0b7ee20564d4d7ca3f263f93e2279

Observation 8685f695-97b5-4632-b1b4-a12fb11e367b · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.570236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:8f76048b176fedd79666400bf08525c633650be6a9fcccf5fe2f2e2505c19d8a

Observation 4584d383-7a8d-4aa6-adec-bfdb1e2a0d79 · inbound

GenVC: Self-Supervised Zero-Shot Voice Conversion cites this paper.

GenVC: Self-Supervised Zero-Shot Voice Conversion Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:27.053709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:27.053709Z digest=sha256:917cd46bd2dc94bfc8b90b966cb12102d8b33be26e25c296807798a832deb453

Observation 0c22a700-ffc6-4926-af3b-64c99eef2eea · inbound

Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance cites this paper.

Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T21:54:08.516605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:54:08.516605Z digest=sha256:547ff37f8f252c10524a6d906b4887813111730b88a3b80d4bd96f102299e278

Observation 240147d3-4c80-4585-b45b-8d478dd37758 · inbound

Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits cites this paper.

Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:50.926162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:50.926162Z digest=sha256:3e2c0b689afe8c6e003a6fef56aef7a4f66be553e94145ef53c2cd95caacdf13

Observation 64fe4a74-6347-4730-9265-6015099d6104 · inbound

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding cites this paper.

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:15.604779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:15.604779Z digest=sha256:1ac02c026981a0da387d1099b0730c06370fb32d9152118803b90ed3497229e2

Observation 7e1872e8-7f8e-468d-bf75-a5acd1a15bc1 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.462342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.462342Z digest=sha256:546e32a104d25f4bd4748696d10f8051242aaff11264e0b42c26e1c9167c8a9c

Observation fab4a743-7c6c-42bb-a1ab-0ef3ae3e293a · inbound

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling cites this paper.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.227783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.227783Z digest=sha256:f463d950e75d93000cb3e29ba3add7725048661d4bfb0853681907b66406153a

Observation f148466c-20d9-4340-a9ed-b9aedfd7a3f6 · inbound

Voice Adaptation for Swiss German cites this paper.

Voice Adaptation for Swiss German Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:07.320866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:07.320866Z digest=sha256:5c262c2d330a76d57d2cf209f8f9f07b020d6287cab617ac267ebaf1cc42a40d

Observation 7ab4c217-426d-4035-a07d-ab117c4c1ffc · inbound

Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes cites this paper.

Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:31.620520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:31.620520Z digest=sha256:8e95bec460720220382fe5478bb9fe7d8069faee1e741b973e3b23233b6c5308

Observation 9a7b921a-32ce-4f59-9765-7ed4dbf538f6 · inbound

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation cites this paper.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.091198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.091198Z digest=sha256:391b969a14908715a2ee5ddfa3e598694a22f44750db734dc2cadfb1ae13c156

Observation 09225bfb-1e78-4efb-bf6e-b9e4b684cb17 · inbound

Kinship in Speech: Leveraging Linguistic Relatedness for Zero-Shot TTS in Indian Languages cites this paper.

Kinship in Speech: Leveraging Linguistic Relatedness for Zero-Shot TTS in Indian Languages Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:30.568112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:58:30.568112Z digest=sha256:b2d3c295d18d40114ae409b9386bf08edb9ac0e9b8348f5bc253a060571aa06e

Observation e58d9ddf-40d0-4d10-893d-ff661ecaf7eb · inbound

OpusLM: A Family of Open Unified Speech Language Models cites this paper.

OpusLM: A Family of Open Unified Speech Language Models Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.334752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:36.334752Z digest=sha256:a138781b55105ac93f52d070077edb11198706c65cb05d41207a3c33209ac364

Observation b87f9175-72f9-4052-ba6c-4c17668cba53 · inbound

Differentiable Reward Optimization for LLM based TTS system cites this paper.

Differentiable Reward Optimization for LLM based TTS system Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:03.927154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:03.927154Z digest=sha256:b93e3304b61f809c431db918b0ced63a970eb06c3b6b369096302671719580c9

Observation 78b9d3b1-68ef-4d18-ae2e-3aabd768eb0f · inbound

SecureSpeech: Prompt-based Speaker and Content Protection cites this paper.

SecureSpeech: Prompt-based Speaker and Content Protection Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:36:07.615777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:36:07.615777Z digest=sha256:aed0dc9ffcbe90bc680d82f9b1f582d4e89d10ed8d5496d679f90a1f402d1692

Observation d4517a79-8173-4f65-8551-ffcd886f9f98 · inbound

Next Tokens Denoising for Speech Synthesis cites this paper.

Next Tokens Denoising for Speech Synthesis Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.950994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.950994Z digest=sha256:ee953d5afb602ebf165073b9e68cfd63f5837dcb93c1a946a6dbfdf74eaaceee

Observation 64a33c59-d966-4a47-b022-18342097d1b9 · inbound

XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation cites this paper.

XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T22:17:38.301601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:17:38.301601Z digest=sha256:3c59ab8e9e584434b81fb0f1e9abfa32f8cfba574673bff575585ad040c4e9e7

Observation d9c9843c-e0b0-4d37-9727-f0540ccd7d32 · inbound

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis cites this paper.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.883671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.883671Z digest=sha256:b572590c6a84beb69f574604c922f28a7e3bfb91a0afec01a2b8cf9775f4accd

Observation 9dc06219-a9e9-4be9-b4e1-e28e44f6aed5 · inbound

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching cites this paper.

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T18:51:21.136655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:51:21.136655Z digest=sha256:955977bf27be3498bb8a1effe70bbcc5b883415f799a4c7c8a50cb9a1bb986dd

Observation c93a56ef-7037-49ee-95c9-33e0a5070a7b · inbound

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance cites this paper.

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:36:28.826038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T14:34:15.220263Z digest=sha256:8d4f9bc272427f5b83c22c43ee421a33481e112bc1f8e62c0e82e326b69ce990

Observation 488ae25c-b8e7-4ae2-b441-d6cc262f33b1 · inbound

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model cites this paper.

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 87

Resolution
malformed identifier
no resolver link, observed 2026-08-03T00:11:18.974337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:11:18.974337Z digest=sha256:289046c17deb84e8313636068dd3fbb3a213ef5c1154e15d7518e763fdb3fcb4

Observation 84b0e2ae-172b-42a0-a79f-bcfdf1f1c229 · inbound

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization cites this paper.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:59:59.597146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:a096028b1674cab9d5dd9c2d8c76e2655e027e7db3b3187d46d91991cda5d611

Observation a5fd5a40-14db-4ffa-b6e1-e1094da1f026 · inbound

StreamMark: A Deep Learning-Based Semi-Fragile Audio Watermarking for Proactive Deepfake Detection cites this paper.

StreamMark: A Deep Learning-Based Semi-Fragile Audio Watermarking for Proactive Deepfake Detection Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:05.181855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:43:16.998414Z digest=sha256:64190d9bdab7a8abf34234e1bca1a1af900e507e5f712f0b7453bc7313936926

Observation 3a70b231-b3aa-409c-87f2-0ca65bf922f5 · inbound

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation cites this paper.

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.473621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:36:07.090627Z digest=sha256:d2845714e4854d53d273cfb50b0b2085834e2acff25a64de8eb7050903165008

Observation 22772b80-26cd-4144-8f72-ad74a903bf89 · inbound

Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations cites this paper.

Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:13.516306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T08:16:47.522489Z digest=sha256:1e3901833d1f9c0a31136845f5f86b40a9f349cc1dd1975cbb1cc3a5de699820

Observation dddbbbb0-4ae5-4e67-9b5d-be997bff803b · inbound

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning cites this paper.

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:16.042449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T04:35:58.032597Z digest=sha256:5c1f085ccdfae95f39314398c6e2010b5bdf6cc5daeee5e9638331c36ce8059e

Observation 8abf31f5-c8c4-4ebf-ba4c-3459173fd0dd · inbound

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning cites this paper.

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:13.915078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T01:43:48.555523Z digest=sha256:1b6b19f4863bd26487b20ac1ac22e4ecaa9eeffe27a80898844df8d5f1a1cdb3

Observation 11430ff3-1019-4358-86c2-e879e4a3d3d3 · inbound

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation cites this paper.

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.333971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T14:08:30.802619Z digest=sha256:2fb1b4ccc368753f62ae296fd2d1473e4fe2f9ad6e414694eab711f1a949bf96

Observation f3ef0fae-32a5-48a0-bb78-c852c98aca91 · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.842971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:1269d9e9e60d8355bd22db634c2f64d7d96fac785aa23a1a36a1d04249fda597

Observation 03a2684a-eeb9-4fbf-8e07-d08c633d1569 · inbound

Do speech foundation models perceive speaker similarity as humans do? cites this paper.

Do speech foundation models perceive speaker similarity as humans do? Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:57:04.871209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T00:06:34.841809Z digest=sha256:d0a14ff8098514838c5e50b31ac87fa18d04dd97eb8cd96e93e9d1e9b84b8df3

Observation b51ec42e-5a70-42bc-bcd1-f49e35a5c523 · inbound

UniVoice: A Unified Model for Speech and Singing Voice Generation cites this paper.

UniVoice: A Unified Model for Speech and Singing Voice Generation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:08.306137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T23:56:14.198308Z digest=sha256:996dad12e8e482b8ab7c785170e9aac9d2410d3dacbf8c0af2de2b3b438521c3

Observation eff54b84-0495-4ced-a6c8-cb7fbd651d5d · inbound

CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations cites this paper.

CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:07.156796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T20:21:47.681999Z digest=sha256:fbe9c40078523c87fbd1265215a7b0e7009994d7121ca9ae5adfc7c1a9a3243c

Observation 23bea1f1-e3f1-4cb7-a49c-5a9264df1192 · inbound

VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation cites this paper.

VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T14:39:57.036118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T03:26:22.529879Z digest=sha256:860e7b84418545f9dfe1b472aa6c76451ef67f36435d1ba1a9ca3c217174f15f

Observation 03b696fa-7ee1-4897-96cd-1ece412df834 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 235

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.394850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:ce327ab077e032a01d0b524b37fca9d7c81dd328fe5ca57b58ab0ae237257af9

Observation e6fa88f4-b57b-4f87-ade2-356a59d72afa · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 235

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:7f059b93ac1b99e5c77c16fb0f46f19bf79accc93ffda90609925aca0dfecc66

Observation 487ed536-8ec9-4116-a1b7-b26693617478 · inbound

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection cites this paper.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:08.442744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:08.442744Z digest=sha256:21030fc874ceb5484e0c78be78525451536eee37201529b2e9705db41568891d