Pith. sign in

Paper Citation Record · LEDGER

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model

As of 11 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2501.05787.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05787 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:10:55.890151Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:27:24.720460Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:27:31.509181Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6d2ad769-a91e-454d-aeb7-05bdc061fd4b · outbound

This paper cites XTTS: a massively multilingual zero-shot text-to-speech model,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model XTTS: a massively multilingual zero-shot text-to-speech model,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.534824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.668281Z digest=sha256:a370f57189b99c030abe43a4f162804cf1fa46de96c8869ce50e6a5a1f64f536

Observation 7634958c-4ce9-46bc-9289-5ceab7bc1b60 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:55.674066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:55.674066Z digest=sha256:05b95de3abe362597b042c30b01c5fe8cb5b59fd69d4e4bf9bbda50a2e2daa20

Observation 612966c6-3ff9-4716-9fec-d46728d57772 · outbound

This paper cites StyleTTS 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model StyleTTS 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.516416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.679953Z digest=sha256:86dc3ae1c6b055faebddec59bf042ff7911a7541dca6d35167187b51baa3fa86

Observation 6cb4f220-8c59-4d00-acfe-63de81674a74 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:55.698238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:55.698238Z digest=sha256:c56993c785ab7c388fcbc91e902ce732f7341539578bee32a5207626de967906

Observation b4f255c2-d14e-4c99-bf3d-df1fbcc3c207 · outbound

This paper cites Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:55.704933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:55.704933Z digest=sha256:42b85ac54456d4d468855667aa778d1728ceff83b31a598ab77814d8f8e10949

Observation ae8cc110-6776-45f3-9fb7-a720f987fc5f · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:55.710896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:55.710896Z digest=sha256:d31e41221dc25df6b99ec38a1cf6766291a81ce4423f4ca6c33456ee8e0d5f63

Observation c0c0cff0-e2d7-42a7-b3cf-88ba5d377703 · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model ORPO: Monolithic Preference Optimization without Reference Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:55.723234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:55.723234Z digest=sha256:be98c7e5fd999dbdaf3486db6e8b08b2b767df7b4572d9a50265c548fc65b90a

Observation 8c3bd006-3060-461b-a99f-ed86fbf2d8eb · outbound

This paper cites EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:55.728621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:55.728621Z digest=sha256:fc6472ac7be4dbf024f34da7cf31c10008333699b705a86c708b5036b8c84ae6

Observation 851a3882-752a-492d-8058-b8c3accc4862 · outbound

This paper cites High Fidelity Neural Audio Compression.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model High Fidelity Neural Audio Compression

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:55.734570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:55.734570Z digest=sha256:7e7a1a09273b2d0449b7b633c21e9c2153ba07f6be871ca9a6da2b774fa1e30a

Observation ce67a96f-ab0b-47c8-956e-3546ac00d645 · outbound

This paper cites High- Fidelity Audio Compression with Improved RVQGAN,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model High- Fidelity Audio Compression with Improved RVQGAN,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.488692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.739930Z digest=sha256:47495ce722f718f098203153a9431fd117f30d64cc30f39b4afc8ff987e021d7

Observation e6f4c04c-85f7-4f7a-a180-a69162e39f0b · outbound

This paper cites Neural codec language models for disentangled and textless voice conversion,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Neural codec language models for disentangled and textless voice conversion,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.472934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.744796Z digest=sha256:4209ac0e572365fce833466bc426dd26110dc38bf40cd4b669d11ad07ca90f97

Observation 103b226a-c1a6-4379-9769-824e366f0f0b · outbound

This paper cites SNAC: Multi-scale neural audio codec,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model SNAC: Multi-scale neural audio codec,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.457008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.751715Z digest=sha256:eb9ee11a85336d77da89816da001cf08f55e94d3c0925eec1250ef5ad3913aca

Observation 11998235-261a-4a26-84fc-ee303e461380 · outbound

This paper cites ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:55.757511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:55.757511Z digest=sha256:290578b11a013ab310ed16ffeba5152886919d165666f383c36f20482a04a0c6

Observation 3c1cf39c-d2e5-4445-a5c5-229a98178104 · outbound

This paper cites LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:55.763819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:55.763819Z digest=sha256:f9b799c3177603eb5c4e2dfa99450b56652f07f6a3fa4112d6ea439d2f4b33ed

Observation 05b3852c-8962-4e07-9feb-03fd0062bf33 · outbound

This paper cites HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:10:55.995809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.769641Z digest=sha256:db4ce93d2e9dcf4a6e0a801a196973d26c0f6943facdbf0672d50f0833935093

Observation 037475ec-231d-4d70-983a-17185034ef21 · outbound

This paper cites MEGABYTE: Predicting million-byte sequences with multiscale transformers,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model MEGABYTE: Predicting million-byte sequences with multiscale transformers,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.441952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.776575Z digest=sha256:a872a398629f073064e2b5b2ae9968b5abd934c6da9d36d7fb24a6bfb0aa0a79

Observation 2f15b612-c94e-4939-946c-ed1c0a5d5daa · outbound

This paper cites Mish: A self regularized non-monotonic activation function,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Mish: A self regularized non-monotonic activation function,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.423102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.783321Z digest=sha256:6f04dc972baa34749292313154bc5cee854f67cc56caa1d2461cdcda69299b13

Observation 6da272d5-64a1-4f6a-96aa-fa9c81cafba2 · outbound

This paper cites Attention is all you need,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Attention is all you need,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.405299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.789138Z digest=sha256:6875bd1ded97ed98bcdb7ff5f12621238c83f6d6de62d5affedd98856fa723b3

Observation 2fa9202a-7895-4210-8688-06245077faf1 · outbound

This paper cites Natural language supervision for general-purpose audio representations,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Natural language supervision for general-purpose audio representations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.386792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.797946Z digest=sha256:dacd905155567a17efa86c259a1644673cb697405277722d3da7fa9d3bf6cf48

Observation 6efa0434-90ba-4650-944f-cacdd33477f6 · outbound

This paper cites A new algorithm for data compression,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model A new algorithm for data compression,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.370717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.809730Z digest=sha256:58191dea245946bc800b247dd3b3d0771d8d7f5476196702e1ba4bc0f68cd48f

Observation e22f5484-237c-4562-a41d-81817995acfe · outbound

This paper cites Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:55.818304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:55.818304Z digest=sha256:90b5da32df5f11d3dcbd50bdf0582edb1a95bc863c74861d63cf71db4bda52d6

Observation acf43f8d-f33e-4f7c-904b-5e359c231b45 · outbound

This paper cites CSTR VCTK Corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model CSTR VCTK Corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.353205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.824155Z digest=sha256:05fed11888807383294d366ebd24d9841ce7ca5cd7185018a4b671a3825ea1d7

Observation c474950c-ab17-43bb-8de9-4ddb3c5cb2df · outbound

This paper cites UTMOS: Utokyo-sarulab system for voicemos challenge 2022,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model UTMOS: Utokyo-sarulab system for voicemos challenge 2022,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.336547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.830053Z digest=sha256:a2e27a152b4e4e5a5b0f8704144fa886ce0ddcb1ea226b694ebac411dcab81d0

Observation 7d2955d8-7dfb-4dda-bc22-9038f8a977f9 · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Autoregressive Speech Synthesis without Vector Quantization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:55.835745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:55.835745Z digest=sha256:f0d2ac345b834492b64e14b7decbc09f22cfa260a572159ee3657497e5f14acc

Observation 90df133b-a885-4aa3-817f-b9452ab21c11 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Robust speech recognition via large-scale weak supervision,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.319561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.840443Z digest=sha256:f4505b64fcc9d31400bdfe00ec3b2a9bd88283189fc887263ad373cc14a2a880

Observation 623f79b8-06f0-4d82-b6ab-13f76155101a · outbound

This paper cites MetaV oice-1B,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model MetaV oice-1B,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.301767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.844964Z digest=sha256:b55fd73cf56acbfded6f68fdc0ba2b5d67be612027bd47b9ffe846039fbe23e1

Observation b38938fc-2d38-4752-aaa2-3f5fb8e48a0e · outbound

This paper cites WavLM-Base-Plus-SV Model,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model WavLM-Base-Plus-SV Model,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.284349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.850062Z digest=sha256:96d48e2a45b8cd970022cb8b19261ddef5f98a8e27e2c4fa8324d2a412dc1cd5

Observation 48f65357-44fe-4088-9d6e-7b716bb8f13f · outbound

This paper cites Decoupled Weight Decay Regularization.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Decoupled Weight Decay Regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:55.855681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:55.855681Z digest=sha256:bafd340fde2966df4a34f0fecc08be4b288c8bf3b4fb6f340f9436e60d49f9a2

Observation 948be4b8-7e47-41f7-9f92-e90505dae647 · outbound

This paper cites Libriheavy: a 50,000 hours asr corpus with punctuation casing and context,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model Libriheavy: a 50,000 hours asr corpus with punctuation casing and context,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.268142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.860339Z digest=sha256:35bdb45bc32cfad15ceab2bb451e4ff00afd9cf2d1afefce49a6137beb5b4cca

Observation 576fb3b7-8260-4631-bbb4-09ef40784e6b · outbound

This paper cites GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:55.865242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:55.865242Z digest=sha256:d555a7c19460cf7ae598a2b63798aa1952acee449fad2401f0146dc57414a6fa

Observation 8ea29d98-d24b-4251-9687-e4a7dedc93cd · outbound

This paper cites AniSpeech Dataset,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model AniSpeech Dataset,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.249229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.870381Z digest=sha256:a96c0e7b71aa1d8e0eee8865c6f7b7305282b2ef2794adde7f5ae82f11d47a82

Observation 562d7012-7897-4b0d-be7b-e1395b8ba8fb · outbound

This paper cites The casual conversations v2 dataset,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model The casual conversations v2 dataset,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.232096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.875304Z digest=sha256:ce19d2e4971fdf53a83f0b489cfde5e67e583e4b6f9b8f63b4434c6dec1af257

Observation 401dcbaf-d827-4549-a161-7b0fdb3a6c09 · outbound

This paper cites X-Vectors: Robust dnn embeddings for speaker recognition,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model X-Vectors: Robust dnn embeddings for speaker recognition,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.215115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.880001Z digest=sha256:b31d2db58b086007b0a268f2aec5517335eec0f5fdc4208ea681becfa885a65d

Observation 1b7c1e97-200c-45b1-b1ef-832d456393c2 · outbound

This paper cites V oice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model V oice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.197660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.884914Z digest=sha256:75f02d8dfb33e3210fb99ef1b8d7efddc7ba8dd9046f424c08c93506fe4fc62a

Observation c06cc304-c1e2-41c1-aded-ab04b16ef3cb · outbound

This paper cites XTTS-v2: A multilingual text-to-speech model,.

MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model XTTS-v2: A multilingual text-to-speech model,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:56.180256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T21:10:55.890151Z digest=sha256:146b6e8308cd80651d2d58049227b4818b0ef3d4c8a5c1ae7ee7105ee2e21407

Pith citing papers

Observation 85cca677-1b93-4b49-a515-98d54bc050b1 · inbound

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation cites this paper.

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:27:31.591486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:27:24.720460Z digest=sha256:fc7c861876147d0012e72acc4ae8ca440158d905508f923870574263d9807ab9