Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:10:55.890151Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2501.05787.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:10:55.890151Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:27:24.720460Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T12:27:31.509181Z
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6d2ad769-a91e-454d-aeb7-05bdc061fd4b · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model XTTS: a massively multilingual zero-shot text-to-speech model,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7634958c-4ce9-46bc-9289-5ceab7bc1b60 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 612966c6-3ff9-4716-9fec-d46728d57772 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model StyleTTS 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6cb4f220-8c59-4d00-acfe-63de81674a74 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4f255c2-d14e-4c99-bf3d-df1fbcc3c207 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae8cc110-6776-45f3-9fb7-a720f987fc5f · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0c0cff0-e2d7-42a7-b3cf-88ba5d377703 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model ORPO: Monolithic Preference Optimization without Reference Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c3bd006-3060-461b-a99f-ed86fbf2d8eb · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 851a3882-752a-492d-8058-b8c3accc4862 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model High Fidelity Neural Audio Compression
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce67a96f-ab0b-47c8-956e-3546ac00d645 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model High- Fidelity Audio Compression with Improved RVQGAN,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e6f4c04c-85f7-4f7a-a180-a69162e39f0b · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Neural codec language models for disentangled and textless voice conversion,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 103b226a-c1a6-4379-9769-824e366f0f0b · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model SNAC: Multi-scale neural audio codec,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 11998235-261a-4a26-84fc-ee303e461380 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c1cf39c-d2e5-4445-a5c5-229a98178104 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b3852c-8962-4e07-9feb-03fd0062bf33 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 037475ec-231d-4d70-983a-17185034ef21 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model MEGABYTE: Predicting million-byte sequences with multiscale transformers,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2f15b612-c94e-4939-946c-ed1c0a5d5daa · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Mish: A self regularized non-monotonic activation function,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6da272d5-64a1-4f6a-96aa-fa9c81cafba2 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Attention is all you need,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2fa9202a-7895-4210-8688-06245077faf1 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Natural language supervision for general-purpose audio representations,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6efa0434-90ba-4650-944f-cacdd33477f6 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model A new algorithm for data compression,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e22f5484-237c-4562-a41d-81817995acfe · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acf43f8d-f33e-4f7c-904b-5e359c231b45 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model CSTR VCTK Corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c474950c-ab17-43bb-8de9-4ddb3c5cb2df · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model UTMOS: Utokyo-sarulab system for voicemos challenge 2022,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7d2955d8-7dfb-4dda-bc22-9038f8a977f9 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Autoregressive Speech Synthesis without Vector Quantization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90df133b-a885-4aa3-817f-b9452ab21c11 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Robust speech recognition via large-scale weak supervision,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 623f79b8-06f0-4d82-b6ab-13f76155101a · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model MetaV oice-1B,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b38938fc-2d38-4752-aaa2-3f5fb8e48a0e · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model WavLM-Base-Plus-SV Model,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 48f65357-44fe-4088-9d6e-7b716bb8f13f · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Decoupled Weight Decay Regularization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 948be4b8-7e47-41f7-9f92-e90505dae647 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Libriheavy: a 50,000 hours asr corpus with punctuation casing and context,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 576fb3b7-8260-4631-bbb4-09ef40784e6b · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea29d98-d24b-4251-9687-e4a7dedc93cd · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model AniSpeech Dataset,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 562d7012-7897-4b0d-be7b-e1395b8ba8fb · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model The casual conversations v2 dataset,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 401dcbaf-d827-4549-a161-7b0fdb3a6c09 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model X-Vectors: Robust dnn embeddings for speaker recognition,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1b7c1e97-200c-45b1-b1ef-832d456393c2 · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model V oice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c06cc304-c1e2-41c1-aded-ab04b16ef3cb · outbound
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model XTTS-v2: A multilingual text-to-speech model,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 85cca677-1b93-4b49-a515-98d54bc050b1 · inbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.