Pith. sign in

Paper Citation Record · LEDGER

VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2305.16107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.16107 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:23:03.314763Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.442888Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c3e9db4e-2e37-4824-b626-ded906c5e4eb · inbound

DASB - Discrete Audio and Speech Benchmark cites this paper.

DASB - Discrete Audio and Speech Benchmark VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:28:39.518355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:8137264f068e7bcd4eca129b58dd7f59c9a0c547e403b756349f39a92174cb1b

Observation b77fe242-bbaf-4c88-81c2-4b7a7701dbcb · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.542344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:fc16a686f04d41f54fa2d14230c2082bd8c7eb1c85fb9de74c19682efcfc682e

Observation 80f70929-f25f-419f-8682-d847b21bd568 · inbound

Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis cites this paper.

Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:23:14.993539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T17:22:23.797551Z digest=sha256:cc08248be7ff0ba8ec64ee455d7188a830ac895d9f08e6037c32a5d6168301d8

Observation 4f65ed92-a07d-4eba-b903-d475e502c3ad · inbound

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models cites this paper.

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:03.314763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:03.314763Z digest=sha256:0ec281c183f64dbfffdaa7983a364bf948a091b48bf07439205850e1cd992f7f

Observation dc3048c9-080b-4860-b527-1fcdc0f87fc5 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.771128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.771128Z digest=sha256:5371c6233e7d74d7467e334c65f1be1b787307388a66616068230e4d0d6a481e

Observation d051f26a-f4d5-4dc1-a1a0-71add2d27203 · inbound

Probing the Robustness Properties of Neural Speech Codecs cites this paper.

Probing the Robustness Properties of Neural Speech Codecs VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:31:37.177571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:31:37.177571Z digest=sha256:241b7cf1d75fa70158cc94d7f4a07c7e302ba7d64b988f9fd0ad99968a4966d4

Observation 4870119f-084f-4438-ab2b-8d42b446748a · inbound

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction cites this paper.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.093509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:41.093509Z digest=sha256:df1072fd5e9ce3150a34e012373ac9947b7ccf710c1b4f4d9c717b8609846351

Observation 82e7c989-1115-41fd-9d0b-10bca00a2b5c · inbound

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs cites this paper.

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T16:46:49.240587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:46:49.240587Z digest=sha256:b0c91383e9b730c6609754a9bbbb1484b473df8e2a1f03a9e521db1fa059afca

Observation 7f3ad701-dcea-4c3b-bad3-7a246672d405 · inbound

Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy cites this paper.

Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:08:08.995130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:08:08.995130Z digest=sha256:68e2d0114033d9c6d3012321e2134b1932b62402c93e2f1730c140765346e6a4

Observation 1716ce80-123e-484e-a4f3-65e6dab0be23 · inbound

The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer cites this paper.

The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:43.836552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:43.836552Z digest=sha256:ae9b8088a961ceaf5371dc3e308e5f3b1919870247735ef57d9877933d3ba4b0

Observation 50646c69-958c-472e-8e22-427f1a6c492d · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.727845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:25:52.847432Z digest=sha256:5f909588b7e1314534cf82ba280f93f74c934fdb2c161d1298d6a21171602635

Observation 7080ea78-843c-43d7-8167-70a5797c4f18 · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:15:07.867939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T23:14:32.494076Z digest=sha256:ebeaf606d51f06ca3d3aebdf5188144d65f4afe02c93e25664dadc2d9d4f964b

Observation 982b8fe7-f2ce-4605-b23c-dcdaf6dfac20 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 237

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.444307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:c1db7f3e910903c2b4fd10a9609ac2ea31b4092244751b393f495036464fc7a6

Observation 095db85d-7326-4c44-b4e9-7aebdf04307b · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 237

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:cf53eedf171f7c6fbead73d9959045281c09922133a81785c073fc40969e3bcd