Pith. sign in

Paper Citation Record · LEDGER

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2507.18119.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18119 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:41:57.488433Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:03:09.632009Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T20:59:47.893961Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eb64e32c-ff90-48d9-b9f2-575e4fb635e2 · outbound

This paper cites On the landscape of spoken language models: A comprehensive survey,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness On the landscape of spoken language models: A comprehensive survey,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.914459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.365862Z digest=sha256:755f3186a7d1eff6fa074156450f68b1500c17b593ac439f06b668e94f16cc94

Observation 68bbeba8-9c7c-467d-b81d-bd99ce2c822f · outbound

This paper cites Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.901450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.371356Z digest=sha256:e1ee13c98cda61614d17d27e052b717dd0701d02a7ee492892e0844fcb26921a

Observation 24041366-a50b-40ad-a380-5a6e1234c62c · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Moshi: a speech-text foundation model for real-time dialogue,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.888581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.376159Z digest=sha256:b8100ea8c007564699df962af25476cb93e34f29543ff194ab83e9612552d56c

Observation 72c5bfa6-bcc3-430b-a2bc-3d0e4a717200 · outbound

This paper cites Speechgpt 2.0-preview,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Speechgpt 2.0-preview,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.875689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.381247Z digest=sha256:71cd4d5149b1debb63c9e7b668b7d11e8e37c13382fb2f587649adada00d2194

Observation 938b6b44-702d-478d-a7f9-d18e29f44c28 · outbound

This paper cites Llama-omni: Seamless speech interaction with large language models,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Llama-omni: Seamless speech interaction with large language models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.862931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.385844Z digest=sha256:7cbcac81782a595e3b1648ec47d8ce55788106500323fc0e272dca1902b08b64

Observation e72106f9-554c-4ef0-8642-7a89c502a69f · outbound

This paper cites Freeze-omni: A smart and low latency speech-to-speech dialogue model with frozen LLM,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Freeze-omni: A smart and low latency speech-to-speech dialogue model with frozen LLM,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.849408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.390739Z digest=sha256:354c5a2831386b60d5eccd7a46d69a27bf8d05e9f70ce3f180b1192dd78b9850

Observation 36bed796-faa9-4a2e-96fc-0c967461f38a · outbound

This paper cites Slam-omni: Timbre-controllable voice interaction system with single-stage training,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Slam-omni: Timbre-controllable voice interaction system with single-stage training,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.836513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.396459Z digest=sha256:ea9340390721581766f5c1868fe63337411fb2b7562facb13a2bbec6cf65f757

Observation 02c73c13-ce25-43d2-8f96-5a2a7c3f9521 · outbound

This paper cites Glm-4-voice: Towards intelligent and human-like end-to-end spoken chatbot,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Glm-4-voice: Towards intelligent and human-like end-to-end spoken chatbot,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.823488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.400457Z digest=sha256:6e2915746dc0bc2f64a9d711917929ef1b61fa4078348ecdd5c34a41fc731985

Observation cc859baf-0b64-426d-a228-f8fba81a3694 · outbound

This paper cites Minicpm-o 2.6: A gpt-4o level mllm for vision, speech, and multimodal live streaming on your phone,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Minicpm-o 2.6: A gpt-4o level mllm for vision, speech, and multimodal live streaming on your phone,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.810776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.404284Z digest=sha256:bd221d05925441adc4a2a26c8868d6acedd30ace2e34f12466a84bda2c8f385a

Observation cdb7173f-a781-48aa-8211-75e898235a4a · outbound

This paper cites Baichuan-omni-1.5 technical report,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Baichuan-omni-1.5 technical report,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.796923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.408621Z digest=sha256:1125ffd600512b94d00e0318648d48056f8db43e3ded6bdbaf5141a783362227

Observation 000ff57d-2b87-4200-8928-eda716400dae · outbound

This paper cites Salmonn-omni: A codec-free LLM for full-duplex speech understanding and generation,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Salmonn-omni: A codec-free LLM for full-duplex speech understanding and generation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.782628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.412961Z digest=sha256:c4b1c508966fd1e73a3375c617eff2481e2f4764e77f1a5aab044025734e8866

Observation d6181bd6-0b70-41ca-a7aa-c952f014718c · outbound

This paper cites Minmo: A multimodal large language model for seamless voice interaction,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Minmo: A multimodal large language model for seamless voice interaction,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.768963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.417459Z digest=sha256:ed44e1868d683d69585ec19bcedd27619908cf97e96eefaf0f2e694b2b2fb57d

Observation 9e6aff3e-aff5-4316-b407-9cb0517125a5 · outbound

This paper cites Qwen2.5-omni technical report,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Qwen2.5-omni technical report,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.755597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.421720Z digest=sha256:174114c2c8b4a3e2f78a3bf6b8c12d6fb770abedafda7b5283a136f58e59bf49

Observation b5cd58e1-1896-4480-91a6-0998d2991aa4 · outbound

This paper cites Step-audio: Unified understanding and generation in intelligent speech interaction,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Step-audio: Unified understanding and generation in intelligent speech interaction,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.742003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.426101Z digest=sha256:1dd490dff6a705504f3d67516b28670f67ed028bcb8b4b36d8cbade1f44fb58f

Observation b5606be8-d984-47f7-be92-d774c481c6c4 · outbound

This paper cites Step-Audio-AQAA: a fully end-to-end expressive large audio language model,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Step-Audio-AQAA: a fully end-to-end expressive large audio language model,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.727510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.430435Z digest=sha256:8417a492b7f0db5c9e6df6a7cd27e9a3c7a5e17f039c2c1ee303c5758c13f436

Observation 7c8a943b-d78f-4531-92a9-a43779cb87a7 · outbound

This paper cites Llama-omni2: Llm-based real-time spoken chatbot with autoregressive streaming speech synthesis,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Llama-omni2: Llm-based real-time spoken chatbot with autoregressive streaming speech synthesis,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.713857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.435030Z digest=sha256:111c5b18f71f51ef1e827a00cbf243920e9eaad40e7fcb92e0756556c6f5b576

Observation f7d4abe7-6ffa-4d74-8890-cef3efe68d10 · outbound

This paper cites Deeptalk: Towards seamless and smart speech interaction with adaptive modality-specific moe,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Deeptalk: Towards seamless and smart speech interaction with adaptive modality-specific moe,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.698773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.439127Z digest=sha256:34095273cb2dd6a3762c1add04af920de0e6c6e51b065c65cfb407b1cce2c105

Observation 16f671c4-e59e-470c-aac9-dc8dc68bbab7 · outbound

This paper cites BoSS: Beyond-semantic speech,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness BoSS: Beyond-semantic speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.685148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.443222Z digest=sha256:3c7d242c5b9b2ab94bf23f9878dfc0994bc5239258ca5019c5c5c495de5df47c

Observation 29a2c081-b8f8-430c-a9cf-a6f7328a828f · outbound

This paper cites V oila: V oice-language foundation models for real-time autonomous interaction and voice roleplay,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness V oila: V oice-language foundation models for real-time autonomous interaction and voice roleplay,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.670291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.447315Z digest=sha256:2c72b9abbbaab4b0aa6fc5b0f1b3823a2972269baac182fd54e499eb79bf854b

Observation 9f8af00c-cc09-4c45-bfea-071d41d386b0 · outbound

This paper cites Kimi-audio technical report,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Kimi-audio technical report,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.657231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.451156Z digest=sha256:c124849f57422106c4cecb39a8c2174c8d2d7285b56e9eef4e6abeddd35f9126

Observation 0243e4ed-6eda-4fc8-b4ab-4ada7e33f86d · outbound

This paper cites GOAT- TTS: llm-based text-to-speech generation optimized via A dual-branch architecture,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness GOAT- TTS: llm-based text-to-speech generation optimized via A dual-branch architecture,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.643553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.455369Z digest=sha256:ed129264c23ef5da7251b67c5a1a59a57b13ad9c06c9ec783b3e348107c9affd

Observation 1bb63ac0-1f4d-412c-a07b-5700418282e0 · outbound

This paper cites TELEV AL: A dynamic benchmark designed for spoken language models in chinese interactive scenarios,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness TELEV AL: A dynamic benchmark designed for spoken language models in chinese interactive scenarios,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.630056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.459831Z digest=sha256:4138827c4bc6c2ea3b294678f0d4dc875d6ab3018afe9c47988cf0b8a67aabf9

Observation 4b62431f-b158-4d8f-b63a-611371451226 · outbound

This paper cites Qwen2-audio technical report,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Qwen2-audio technical report,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.616374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.464545Z digest=sha256:3ffa5511bcd519d4f2d0a300c9e7280fa428f9122e904d1b919128338ee860a2

Observation 4385ee2b-2e76-42bf-8e1e-6c072eb367d2 · outbound

This paper cites Baichuan-Audio: A unified framework for end-to-end speech interaction,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Baichuan-Audio: A unified framework for end-to-end speech interaction,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.601830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.468284Z digest=sha256:05937264a04a56ddb50f156fa6f1ca9bf23759c10d65954e576748962e599826

Observation b846dc80-2c58-42c1-adda-536d03c04710 · outbound

This paper cites Telechat technical report,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Telechat technical report,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.586620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.472468Z digest=sha256:c7f00ed4646aa211ee004009f480fd1738178b1ca541e0bd393cd52f14173f2f

Observation a872b40c-1ecb-4efb-8603-fad8c29bbf85 · outbound

This paper cites Audiochatllama: Towards general-purpose speech abilities for llms,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Audiochatllama: Towards general-purpose speech abilities for llms,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.572007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.476480Z digest=sha256:d1f0c3c1586303782796454c3e9d17a38e91db6b57c4380821657041ada5ad5e

Observation f5806a61-abdf-41cb-b52c-e4bb51384651 · outbound

This paper cites BLSP: bootstrapping language-speech pre-training via behavior alignment of continuation writing,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness BLSP: bootstrapping language-speech pre-training via behavior alignment of continuation writing,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.557553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.480339Z digest=sha256:ce3dd0175b3c9f16a459808dea468a40a1e5e7514efebcce3962264ce6d1d994

Observation 6db03b0e-9444-4583-a2ec-65646a9a624d · outbound

This paper cites Wav2prompt: End-to-end speech prompt learning and task-based fine-tuning for text-based llms,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness Wav2prompt: End-to-end speech prompt learning and task-based fine-tuning for text-based llms,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.542387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.484253Z digest=sha256:32cf7e66f23d69625f7b0ba6f3e8857bc2297b8c20b8772d52c43e10fa829538

Observation 3bac4abc-912c-46b0-87bb-31a4d20a8f0f · outbound

This paper cites DeSTA2: Developing instruction-following speech language model without speech instruction-tuning data,.

GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness DeSTA2: Developing instruction-following speech language model without speech instruction-tuning data,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:41:57.527293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:41:57.488433Z digest=sha256:b58bd7c1b33e5e34e042d32deaeba534e2b2323a36f7870d952524493286ede4

Pith citing papers

Observation 114b1b5d-ca08-4e4f-99be-214cde165239 · inbound

BridgeTA: Bridging the Representation Gap in Knowledge Distillation via Teacher Assistant for Bird's Eye View Map Segmentation cites this paper.

BridgeTA: Bridging the Representation Gap in Knowledge Distillation via Teacher Assistant for Bird's Eye View Map Segmentation GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness

Reference 2008

Resolution
verified exact
local_arxiv, observed 2026-08-05T20:59:47.953809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T20:59:43.785209Z digest=sha256:d9f39bb7660cf22809817ed4fc5b26d638a1fdff26d3483c58f976e457ececec

Observation 5affa23d-c094-4c16-896c-d006ce8796cc · inbound

OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue cites this paper.

OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-05T21:03:09.632009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:03:09.632009Z digest=sha256:0d1430635c7e4f07ed7fec1a262afd3231986299fc36458cd0dd7ac675332d8d