Pith. sign in

Paper Citation Record · LEDGER

Voxtral TTS

As of 14 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2603.25551.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.25551 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T21:18:22.911332Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T19:47:19.722749Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact6
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch12

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e099125-5d51-42d8-9179-f8b163ee6863 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Voxtral TTS Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:6ed5d9b1538591966bf6cb2194766a6066f034811348d016f1ff63a8d592a99d

Observation 93308bf7-4b84-480f-88b4-8949f5fdbb8a · outbound

This paper cites Chao, W.-H.

Voxtral TTS Chao, W.-H

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.572678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:1d4fad6da7cf19f75047fd9dd5f37a7fc118a30372cd7e618b4c3284d049cd98

Observation e5aa4e0a-d80a-4e22-ac8b-8a9a02fcc81c · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing doi:10.1109/TASLP.2023.3288409.

Voxtral TTS IEEE/ACM Transactions on Audio, Speech, and Language Processing doi:10.1109/TASLP.2023.3288409

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.576135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:b692615b46cb727e42af05e36f64c513e4361d733dc0e622cc6026d83b096ecd

Observation 009e605d-86d5-48b8-90b3-39a3e07fa395 · outbound

This paper cites High Fidelity Neural Audio Compression.

Voxtral TTS High Fidelity Neural Audio Compression

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:39:35.889040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:a5d0fb6a60a8f243340fb52091d04706e0d07aa0edf1b97962bb0cf04428cb18

Observation 8b662cfa-d9e5-4c14-bfb7-c231a98b59ec · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Voxtral TTS Moshi: a speech-text foundation model for real-time dialogue

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:39:35.893639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:89316e8031258948450180b0b27d0a50f1ba3a10020828ded78911290bbad066

Observation 33eec127-4b32-4e78-be6d-f7889fb1f727 · outbound

This paper cites ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification.

Voxtral TTS ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T00:39:36.281758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:f6f93ee7264d4bce67ce8cb48c84c9a658d3765af5c9951bda3bc85352c021c4

Observation a2ea90cb-c396-4a88-ac44-bda34d8268de · outbound

This paper cites Classifier-Free Diffusion Guidance.

Voxtral TTS Classifier-Free Diffusion Guidance

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T00:39:35.561776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:0153260158563aedeff78573dbcbfb8ebcc308abb56da4f09c9c4e426fcfca3b

Observation fe942684-64d2-445d-8bd2-16a0564a019f · outbound

This paper cites Classifier-Free Diffusion Guidance.

Voxtral TTS Classifier-Free Diffusion Guidance

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T00:39:35.889997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:46cbac1c18380deda0056db59e9aa259333383ae22678de90dbb6a6e294c9abf

Observation f3f11dd4-4422-4421-82ab-e01f5fd3a5a2 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =.

Voxtral TTS Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.565964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:ed36191d59c926c0a457dfc7f199300d83a0299e164e1c51ebd55a6b07dbcf3a

Observation b181bd6f-437d-4329-a365-cbee707a207d · outbound

This paper cites Alexander H Liu, Sung-Lin Yeh, and James R Glass.

Voxtral TTS Alexander H Liu, Sung-Lin Yeh, and James R Glass

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T00:39:36.279560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:f67e5c43d14ef20397e96d0ab6b69cee1b5f5532fe4efbcbea7497ede5d2d78b

Observation 1ed46d8e-0b15-47c3-a503-5d2f63050ffb · outbound

This paper cites Voxtral.

Voxtral TTS Voxtral

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.904186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:f299ed1be4048fd2d617bb071c37b5b34717f2c8efb300ec3f524f1bd815bc19

Observation d8ea44ac-40cb-4828-8173-552871608972 · outbound

This paper cites EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis.

Voxtral TTS EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.896911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:639a1833ae62bce6bd05ad213780a70f7788e3689bb5872968a4007bbb955cf1

Observation ec8e546d-e712-4730-9e1f-5c7fb70d555c · outbound

This paper cites Scaling Transformers for Low-Bitrate High-Quality Speech Coding.

Voxtral TTS Scaling Transformers for Low-Bitrate High-Quality Speech Coding

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.910557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:2864e562d43b9ceec41b4e1c4c6bd6e551804930a643fe2187e87a08c2c46618

Observation b16f17ab-2d64-4dd7-b86e-2efe76bc2476 · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

Voxtral TTS Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:39:35.906928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:815c97cd83974b37046f56d0b47aaff0a2084b98bc7b2048fc3da794c2bbfc20

Observation 33657dbe-719b-47bd-b6b8-78e3e3089ee3 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Voxtral TTS Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T00:39:35.899779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:38adc649dcd5956a22660c5aaae7c4a672f717d89b17f02396cc6d98b24f4ce7

Observation e003bb2c-f50a-4ffc-bc02-cab6c35ad609 · outbound

This paper cites STAB: Speech Tokenizer Assessment Benchmark.

Voxtral TTS STAB: Speech Tokenizer Assessment Benchmark

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:39:35.914011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:0ccfb39984a98e9c5be61e47e87d774c701bc6dcb2b9ab6227c9d94630228774

Observation 59891d00-1594-4f9d-ba16-054ea3478155 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Voxtral TTS Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T00:39:35.900889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:e1efb79fa80c5c55bb0df5a304d70b3f104f26c0fb73a6a170534375ff3678fc

Observation 1577ae56-63b0-4d02-a211-014364a9fed7 · outbound

This paper cites vllm-omni: Fully disaggregated serving for any-to-any multimodal models.

Voxtral TTS vllm-omni: Fully disaggregated serving for any-to-any multimodal models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:39:35.892375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:8b3534a7037f6b29590b08c236c2c00293e70ffe053bcd2d0818c2bef35bfc30

Observation 0337724a-9c59-481c-9ea8-1557a43548e9 · outbound

This paper cites MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder.

Voxtral TTS MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.876040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:0504d94628db08e4d1b1a5fc1544769fabbf5a103a9be23d006e570cdb0dbfef

Observation cec73bd5-6426-4c71-9705-fc42a7a5ad2c · outbound

This paper cites arXiv preprint arXiv:2512.10264 (2025) 4, 5, 6.

Voxtral TTS arXiv preprint arXiv:2512.10264 (2025) 4, 5, 6

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:39:35.872905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:8a3e016513ec60f4092e4e4d8e228668d1cada817a3eee48f4e952b220cdb7c8

Pith citing papers

Observation a25acf22-ec9e-41c3-b622-2d1b39e2fb65 · inbound

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech cites this paper.

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech Voxtral TTS

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-21T02:33:55.495102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T02:32:26.122526Z digest=sha256:52283f318c539f91e8df0efe81a7aac1026e1180843861ddb27d6cbaa65334f1

Observation 6051f762-2d09-4799-a60a-d30249337f90 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report Voxtral TTS

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:47:19.724003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:5d97c826db273cf4521a1b202f6976b551801111b26c3ae24d825dd847bf2204