Pith. sign in

Paper Citation Record · LEDGER

VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2305.16107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.16107 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:41:45.414147Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.442888Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c3e9db4e-2e37-4824-b626-ded906c5e4eb · inbound

DASB - Discrete Audio and Speech Benchmark cites this paper.

DASB - Discrete Audio and Speech Benchmark VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:28:39.518355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:4e0ff2974d1a1d9579ef12bbaffa28a7de54dc83b496c401bf88836f127f6d71

Observation b77fe242-bbaf-4c88-81c2-4b7a7701dbcb · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.542344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:7783daee14962a14eafcf4a665c1d50283c2c6fdab53721f3395d3bf1b35edb0

Observation 80f70929-f25f-419f-8682-d847b21bd568 · inbound

Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis cites this paper.

Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:23:14.993539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T17:22:23.797551Z digest=sha256:dfe47df055599c9198dceab9841e84d13af64d2d4dc9ac3b5a713b029e946d4e

Observation 1a76d6cc-4751-466a-8922-2d176bc48e85 · inbound

Speech Separation using Neural Audio Codecs with Embedding Loss cites this paper.

Speech Separation using Neural Audio Codecs with Embedding Loss VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:41:45.414147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:41:45.414147Z digest=sha256:2e0512eb396013e253cf726bc0ad41ac85ac18d182af546b7ae4d1209ea2ce52

Observation 84bc28e7-5e5e-4a42-803e-08556780eeed · inbound

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey cites this paper.

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-10T14:36:19.696114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:36:19.696114Z digest=sha256:46648a4791830f2d1a07723afcd61e63a7948eb51b4f7f9d8dbbee8157e5e1f7

Observation 4f65ed92-a07d-4eba-b903-d475e502c3ad · inbound

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models cites this paper.

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:03.314763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:03.314763Z digest=sha256:56570116e7019445a0f0aa17c4bc900e144e30cd3f494f05aef7bee404426634

Observation dc3048c9-080b-4860-b527-1fcdc0f87fc5 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.771128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.771128Z digest=sha256:96159a044cd849e310787829e70c534d9da5b144b573815af37e6a7ba195faf8

Observation d051f26a-f4d5-4dc1-a1a0-71add2d27203 · inbound

Probing the Robustness Properties of Neural Speech Codecs cites this paper.

Probing the Robustness Properties of Neural Speech Codecs VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:31:37.177571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:31:37.177571Z digest=sha256:241b7cf1d75fa70158cc94d7f4a07c7e302ba7d64b988f9fd0ad99968a4966d4

Observation 4870119f-084f-4438-ab2b-8d42b446748a · inbound

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction cites this paper.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.093509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:41.093509Z digest=sha256:03ebbe672b05a736841c1008a42d6cc1345846b98ac5f28eb8d5947330d0ed5e

Observation 82e7c989-1115-41fd-9d0b-10bca00a2b5c · inbound

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs cites this paper.

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T16:46:49.240587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:46:49.240587Z digest=sha256:dc798e2e230662c3e128104780cffb05cc2869afe455e0794fca0c3a3c6e24a8

Observation 7f3ad701-dcea-4c3b-bad3-7a246672d405 · inbound

Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy cites this paper.

Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:08:08.995130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:08:08.995130Z digest=sha256:68e2d0114033d9c6d3012321e2134b1932b62402c93e2f1730c140765346e6a4

Observation 1716ce80-123e-484e-a4f3-65e6dab0be23 · inbound

The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer cites this paper.

The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:43.836552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:43.836552Z digest=sha256:84a2b107703023b85e9d9e98090db026d43ac422aeee02ebaa014630ee3a0fb4

Observation 50646c69-958c-472e-8e22-427f1a6c492d · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.727845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T12:25:52.847432Z digest=sha256:37284de370be88694106b070bd86f2ec62a995fe28856e5243236c7208b77518

Observation 7080ea78-843c-43d7-8167-70a5797c4f18 · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:15:07.867939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T23:14:32.494076Z digest=sha256:29be270a1f02f7468890c3084391735d4d1729577f273892bb9a3d7487661b68

Observation 982b8fe7-f2ce-4605-b23c-dcdaf6dfac20 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 237

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.444307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:08bd8a90a78caf934676e0cf673fdf7abb4174f471443b1bbc0212c86f3a94c1

Observation 095db85d-7326-4c44-b4e9-7aebdf04307b · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 237

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:11970232552fa1788b5711a0cebc5724e4618f284297006212c435b4e6831809