Pith. sign in

Paper Citation Record · LEDGER

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

As of 14 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 10 inbound Pith citation observations for arXiv:2412.15649.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15649 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:18:40.522869Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:32:24.029868Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:18:12.775752Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ddbcb98b-f99e-491f-850d-c25103875095 · outbound

This paper cites online" 'onlinestring :=.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.470848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.470848Z digest=sha256:840490b79e6da3dc15b83f563a4904c495efe1e10511b36c463bd9893ea942b1

Observation e83085ff-21c2-441f-8abd-09ae069ae99e · outbound

This paper cites write newline.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.548428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.548428Z digest=sha256:1cc1897f8c2f33d9e7e2faa5b38c0338257e53b70306bb23ff8ee8b78907d1e8

Observation 0243ff5e-c9e0-4320-a928-865ebea1d1ca · outbound

This paper cites GPT-4 Technical Report.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.591288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.591288Z digest=sha256:34cb67afe3e495a9614721d263459e6c36c85d3c95e316fe4b648de149a9fa5a

Observation 6646d445-79b5-4725-b853-f2045132588b · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.629693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.629693Z digest=sha256:66456f0cf441b9f12ba4c5c03fbd887f90c3a29c54843ac0f162b5f05cacfadb

Observation e9ca6d90-c988-43e9-8243-f2763bb735e8 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.633711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.633711Z digest=sha256:496c12633d6156b5b139b639149c94a69b59a6b7c3d8d781bce4f8ffe3974355

Observation 9bedda1e-26fe-4ac7-9716-f8d7f61661de · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Common Voice: A Massively-Multilingual Speech Corpus

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.639543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.639543Z digest=sha256:cdbde9011048e3a2c6bb0ea476f5b155e549fb4f30c8f742c45502dfd908f042

Observation fcb239eb-9466-422a-aa2d-3a257d5f1400 · outbound

This paper cites MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.644586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.644586Z digest=sha256:7a6f9274b85bad7424f37043628e2dafb4de56ead39ab29ffa1d585c2b086f05

Observation 7a12acbc-2928-4026-8100-ffad69bab4e1 · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:18:41.335383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T11:18:39.753307Z digest=sha256:0485881287588f78a2c3bae1f9139bc5d7e7cb3bccfdf84bea6ebb16fabe9c14

Observation da4b9510-7d37-4deb-ab38-3e54ffa94035 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.782229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.782229Z digest=sha256:34d618ed2eb9ec46d04b2c9c3062174f8dd4ac10543ce585b9895fb2c0f55453

Observation 9c99d1c8-2276-4e12-bead-355ff292e314 · outbound

This paper cites VoiceBench: Benchmarking LLM-Based Voice Assistants.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training VoiceBench: Benchmarking LLM-Based Voice Assistants

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.785833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.785833Z digest=sha256:dd0a06dd50021ed041eb701c0c5603b2998b6fe13bd2b31013b884ba78da1da5

Observation 8bed62e0-65e2-46e4-944f-dbcdc3f73cc3 · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.789782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.789782Z digest=sha256:3b26be696c87fd2caa6389d57b2493bfd8e594c6808a3629efe3c992516714bb

Observation b733b015-4871-4e6b-9d17-0a4ce6535436 · outbound

This paper cites High Fidelity Neural Audio Compression.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training High Fidelity Neural Audio Compression

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.794488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.794488Z digest=sha256:867ab3846e81176dddcf2f91524d5e579606a8bbee4facd2acb87773d5f788ff

Observation 32a30e75-cec8-4151-8087-a7c8f5fe1346 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Moshi: a speech-text foundation model for real-time dialogue

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.851181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.851181Z digest=sha256:e0190ba2ff557c84741daacbf1aae971f4e0a007ccc31155cdd11e8268be72d5

Observation d130f468-0397-46b1-ae1c-316c7b770213 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.902223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.902223Z digest=sha256:270fbaf25868393efa7128261a187239620f7709230238e289d4e56c6c402aff

Observation e9e7e71d-b9aa-4077-b5a0-9a430e071d00 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.907370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.907370Z digest=sha256:43d782725527a951add4c7d175b7c4e86203eb89a89f4c65c02df071fc0dcf68

Observation 68e1af96-30e5-418c-9825-7b78dd791802 · outbound

This paper cites The Llama 3 Herd of Models.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.911264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.911264Z digest=sha256:03e29070c3562ae316306261a0e8d964efa45385d3985983c2739a8c61fed4fc

Observation a6b9e071-c877-45b1-afdd-4bd4948362f2 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.915261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.915261Z digest=sha256:9752c3cb41c43fd98a8d4b5a3f4a330e28ab3798af990c5219c27d907876175a

Observation 4382f760-a711-4ee5-b5ae-8b9d71ea80c7 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.919302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.919302Z digest=sha256:ca990944f1010759e69d9490e8a64c3480711401c46330cab3bb86544dd698f2

Observation 39366192-dac6-481f-9ab6-facd55d3f5e5 · outbound

This paper cites A Corpus for Understanding and Generating Moral Stories.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training A Corpus for Understanding and Generating Moral Stories

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.923266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.923266Z digest=sha256:11ab79ecb364aed7e288f9cf51dec638f491be5115f5e5c0b4af1b505c72f129

Observation dbc72e4f-adc7-45e5-a90a-d006e925a7e3 · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.927221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.927221Z digest=sha256:c43f48c3b1aefe64721c8e3b88dae9d7b8e3371e27b9e010f6551073e67036be

Observation 09f4424b-ffcd-4ce6-a9b4-1e55d6eb61dc · outbound

This paper cites LCSTS: A Large Scale Chinese Short Text Summarization Dataset.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training LCSTS: A Large Scale Chinese Short Text Summarization Dataset

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:18:40.971085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T11:18:39.930688Z digest=sha256:cb4b1b372e48f66a80407a9704269a3449ee6b64b5f47734155ffdece77c051b

Observation 85f77c63-dc55-4ddb-942b-4ab0570238d1 · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training WavChat: A Survey of Spoken Dialogue Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.935097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.935097Z digest=sha256:e8459454f60be25f4ce2ad6c9bc2b4066044483813328f4c26e73262267254db

Observation 82945976-e661-4e4e-9b2f-ec765e720055 · outbound

This paper cites Exploring the Impact of Instruction Data Scaling on Large Language Models: An Empirical Study on Real-World Use Cases.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Exploring the Impact of Instruction Data Scaling on Large Language Models: An Empirical Study on Real-World Use Cases

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.939255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.939255Z digest=sha256:f4d788b63d016725d20b50872cd7260ab268bf4991f583fc5c79a67196a49f6a

Observation f8ebe8fd-1b2e-48bd-a50a-f31102c79e6e · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.945789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.945789Z digest=sha256:76abf42d7411b20a94219d57d944c956853be8c086ff461b77edb86b62ab15e2

Observation 330b4b18-429a-430f-b5f6-5bb0986cc305 · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:18:41.279260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T11:18:39.955831Z digest=sha256:ffd8f4e0ad920b479bd4754ce6fd2ff7ea934113da9c7974e78f737b66398d5c

Observation abce87c4-abe4-4d13-8ab5-a161a7012173 · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.959731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.959731Z digest=sha256:bb1403d488ede5a4498c2de9cf7e7ec7ae9db0808cd9f2b3ef09000907a8b46a

Observation 6f8872c4-908a-4ffe-bd73-8a315dda5b4b · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.012706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.012706Z digest=sha256:a2f9c9efef1e1c76ee3856d3b4c9513aabc78026b18cca6b7a0f417a49e3c67a

Observation 141e6d54-130f-4f56-bd88-f1b2aac659e1 · outbound

This paper cites Decoupled Weight Decay Regularization.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Decoupled Weight Decay Regularization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.080442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.080442Z digest=sha256:1e69e81958483f9224460769a8cd1ef3854ddeff77d1a3a228ab9fa4ec8eaa62

Observation 7dd56083-4d5d-40e5-8eee-4edecf9f44c6 · outbound

This paper cites Language Model Can Listen While Speaking.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Language Model Can Listen While Speaking

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.165779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.165779Z digest=sha256:09a6b6f191f7d5e4fbce7d39334277f2461c35fa0f9f4fa9f62a574a262b49d3

Observation 754d09cb-3d67-460c-8aae-a341229697ad · outbound

This paper cites An Embarrassingly Simple Approach for LLM with Strong ASR Capacity.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training An Embarrassingly Simple Approach for LLM with Strong ASR Capacity

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.272578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.272578Z digest=sha256:bd5b217047e13104867eb043eb56324542c20285db00fb9654904fe292e63c86

Observation 897e0a9c-4bec-447e-a42a-601ff5a3a95c · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.276761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.276761Z digest=sha256:d691d473ee2ee31a485eaaff527e8414b1ef424fdca372ec0487dbe6aa8224b0

Observation f56e587c-99c2-48ed-b502-d968d05b5375 · outbound

This paper cites PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.281167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.281167Z digest=sha256:64f0af1bf424574e70dc2b35fd565df5d866fd5618fa60696fe25ee08f91c7ab

Observation 3cb85342-c33a-43e9-8a76-931dcba67608 · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Spirit LM: Interleaved Spoken and Written Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.285303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.285303Z digest=sha256:9a8db3052bb5239ccd5f4866cf3d31347dfaff9ec53a40e93bdd73483c43dd5b

Observation 39fba148-f830-467e-bfa5-0e8ddbb82c74 · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:18:41.244889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T11:18:40.289562Z digest=sha256:2666bf19f7ba93051d2de4a2cbd425e0d1630650a4dc420ea6ff0097c2068697

Observation 7d62b5d5-6219-4b35-a24d-f4c26ac9e0ff · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:18:41.230750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T11:18:40.292799Z digest=sha256:234ab34fd53bc12ff50ddc775ed9d12639ce9cb8bbded6eabbb124281b562e03

Observation de3a80aa-6cc6-4dfc-ba39-3a686575559b · outbound

This paper cites VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.296375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.296375Z digest=sha256:9411ce0e09d6e7616bf447a47b458599155275fb9d4f4ecd08e409160837250f

Observation b6369a44-f92b-422e-86fe-2bc974adc059 · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.299843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.299843Z digest=sha256:8b81213cfd29336332f7c68340307e577b114c55eb53c911a2c378c522652222

Observation caac40a2-f337-4233-9396-d9a0b5b9b09a · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.303273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.303273Z digest=sha256:f4e451408f407ee7d21c61681398e317aff9416c650739fdf3a61b22e877d476

Observation 6e1924e2-d598-4bf7-9d26-8bc01c011287 · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.306659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.306659Z digest=sha256:fa94dbe2de9a547e07a570d071a46d9405b24b7ea4d065b759bc5b8e6a9273b8

Observation f5a3ae68-a2de-4725-b868-0f426f862264 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.310427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.310427Z digest=sha256:497559c1d5a1daa36ba1b190a42a327190a1f250157cd9e97577e3221d111e07

Observation 0438a759-7569-4c47-b734-ffebbdc8970b · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.313983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.313983Z digest=sha256:2fb05db9f9c8ac08d17a75e38bb88f04a72d43c473851541a7ff057f36bbc3d8

Observation 7c824507-10c2-43dc-9a6c-b270c09b4131 · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.343120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.343120Z digest=sha256:48bd90fe9a43af2ba322a1a16e4dede90abd094b387e37a739235383a11134d2

Observation beded3b0-8a18-4c26-abc3-c6c025291f37 · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.421279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.421279Z digest=sha256:e5673bd845c8dff1665aa29a5f7d2d4614b0658be4a59e4dad264356fafa94a0

Observation 3955d7b3-f724-43ad-a180-de007bd56207 · outbound

This paper cites Qwen2 Technical Report.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Qwen2 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.447005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.447005Z digest=sha256:4d05fc1f9f5481605cd05a8ebf0693513d7254e4f809ca55e77467220704c7d3

Observation 6eeafa04-2715-4596-bc18-ff1178ad1d7d · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.451045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.451045Z digest=sha256:a99f31b5404779408820a2ce416c9025db610fdb52fbfbc699cdca5167626b8b

Observation 57dd6b34-a653-4f7f-bfaa-bb9334055bbe · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.455391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.455391Z digest=sha256:e6482c31b165fc70bb3fce50335c05fa14167fc8e10379aad6b3752264dc3fa8

Observation 81f76123-dd3a-464e-91bc-16ec47d7dc7c · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.459612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.459612Z digest=sha256:31573170b94652d33f2677900bf354cecb8cfeba08037767e3ba562d07a59ff0

Observation a3f8f60c-e12e-49db-9c4a-0b24462bd9c6 · outbound

This paper cites Scaling Speech-Text Pre-training with Synthetic Interleaved Data.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Scaling Speech-Text Pre-training with Synthetic Interleaved Data

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.463008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.463008Z digest=sha256:d47a496e35aa200bf47409722f239d7cdf4154a28b83eda8c0e0cbc87adb6aa2

Observation e9e83824-f1c5-420a-93dc-a24743da2b36 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.467445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.467445Z digest=sha256:a42dcf5574ba1455b4e8de53f013185bf74eacda6633af2f3961590fbfc535d0

Observation 4128e765-f20e-41c7-abfa-7ff1187d36ef · outbound

This paper cites OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.471370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.471370Z digest=sha256:b800b8c9731490182e1a94168e5eb9fca2ac688b90d024dc1901258c61dbde05

Observation 3d875b17-fcfd-4be9-be9d-05520b2506d5 · outbound

This paper cites WildChat: 1M ChatGPT Interaction Logs in the Wild.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training WildChat: 1M ChatGPT Interaction Logs in the Wild

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.522869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.522869Z digest=sha256:1d5cab46080ff34ac864a209759010f86a1443b5093fdcbb870f276c57a9ed84

Pith citing papers

Observation 8a279d26-5421-4784-a928-f7085b96354b · inbound

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey cites this paper.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.029868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.029868Z digest=sha256:a554d130dc9a17cc4a7bb61935a0926b2e00c27be9573cb482232c38a3419746

Observation e4443bba-4982-4deb-b57b-e7ec85031f24 · inbound

Real-Time Textless Dialogue Generation cites this paper.

Real-Time Textless Dialogue Generation SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:27:45.525857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:27:45.525857Z digest=sha256:1674ef79475f7bf6e488d69075fa0df4485e90b8a033996e2c4b2dd9f516c8f0

Observation a214485f-e063-4ae5-80fa-d8217ce3ef7a · inbound

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her cites this paper.

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T13:34:25.641434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:34:25.641434Z digest=sha256:f80de536a40e9c6f5be16cbd64eed79a8bc674a068af7bc0fe7890d8aa6ea3be

Observation 370d4204-988e-4b50-a439-f59ff1b8a613 · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:45:08.035913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:d0aaa416292ebb7c9682c323b3503ce22ac9ead6b8f8b48c0c16fd8a4afecac6

Observation 20ba2d0d-0fac-499b-a766-14073c675d5b · inbound

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model cites this paper.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.356530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.356530Z digest=sha256:2a17d4d8c3fd81c6edc2fded161f495ac4737772b01f097b2ffccc444899ec8a

Observation 81f83c4b-45f8-43b8-9d37-016b0153cecf · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.772401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.772401Z digest=sha256:30b689ded49bfce9742d6f4e1d05c217d9fcb50b7613b4ebb240d2e2d1a7615e

Observation 619f1271-d061-4ea1-9907-7f6fc93723ca · inbound

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning cites this paper.

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:25:30.001718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T14:22:25.660785Z digest=sha256:aa0ea01ece07903fbcaf40197e04be0bc79de4fc4455910b45789583a4029f2c

Observation 9cb079e0-7989-4bc5-bb9d-7dc195febb24 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.237941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:1e071ad44ff9f4a425662008e5b1539584f4ed8a90c970171737eaf42b2a4a01

Observation babbd305-f199-4614-9712-658ef4f26b97 · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:18:12.777326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:c7d602b86a611d0852539452cd3e35cc42470d69713fcee08ed433299457e356

Observation 46fb2046-e84d-4c07-840c-83b750e7cd33 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:08.148931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:08.148931Z digest=sha256:ae6b6941c0c1c3ac70747f9dc326167d833de39d95237bea504ab1d1e8f43286