Pith. sign in

Paper Citation Record · LEDGER

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2412.15649.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15649 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:33:19.356530Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:18:12.775752Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 370d4204-988e-4b50-a439-f59ff1b8a613 · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:45:08.035913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:4a570c3c0a1dfd987b9e0aab859faa94d6fde6577d23004439e3e8248f672f10

Observation 20ba2d0d-0fac-499b-a766-14073c675d5b · inbound

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model cites this paper.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.356530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.356530Z digest=sha256:7a00f1ea25cd12dd439888f93878e44fcf73620b1958d3844b0aec30446d61ac

Observation 81f83c4b-45f8-43b8-9d37-016b0153cecf · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.772401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.772401Z digest=sha256:e3ad98c3410117218f37383459fbd92fcdd37298b1f12d120d73c1eac4c42173

Observation 619f1271-d061-4ea1-9907-7f6fc93723ca · inbound

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning cites this paper.

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:25:30.001718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:22:25.660785Z digest=sha256:d8e063e69a8e2fdfe8f3e5c143e2850932b7fbe074b2481f3e78a36299848bc8

Observation 9cb079e0-7989-4bc5-bb9d-7dc195febb24 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.237941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:18b81baef8db19b75fb79357009e8de692b059d4dea636cec08214383b496c25

Observation babbd305-f199-4614-9712-658ef4f26b97 · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:18:12.777326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:02a239e161b4d95e48f97ac0295c5bd85015ea7b12b461403aaab0f75a47db8c

Observation 46fb2046-e84d-4c07-840c-83b750e7cd33 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:08.148931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:08.148931Z digest=sha256:78a2329dc3e83a8943eb40b75da9d32b88fc875958702137d4993c1320a40f09