Pith. sign in

Paper Citation Record · LEDGER

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation

As of 19 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2505.13338.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13338 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:17:33.475816Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:17:33.349929Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T20:17:33.661042Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation df2a375e-9c0e-41e4-8847-09a411b397de · outbound

This paper cites Recent speech-LLMs, such as GPT-4 [1], Qwen-audio [2, 3], SALMONN [4], and MERaLiON-AudioLLM [5, 6], have demonstrated remarkable performance in handling speech-based tasks.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Recent speech-LLMs, such as GPT-4 [1], Qwen-audio [2, 3], SALMONN [4], and MERaLiON-AudioLLM [5, 6], have demonstrated remarkable performance in handling speech-based tasks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.932117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.345621Z digest=sha256:8a629b812b2891690ae91d92829362bc379dc7c0703a0d5305ccbe35b2965b0a

Observation 5b3e0b13-e110-4c94-b1d6-097889e13798 · outbound

This paper cites Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:17:33.666296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.349929Z digest=sha256:c81d416c1e984852fd97c7b0133f052b89fdd88a2f1df8078e57c47e69d99767

Observation 44a15b1d-4201-4976-9b78-db81f5b07cdf · outbound

This paper cites What is the content in the audio from the text transcript?.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation What is the content in the audio from the text transcript?

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.921636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.353733Z digest=sha256:b89d68f6053aa450f419b9422fe24a8245bcc416cb9ef013edd3db74ef097752

Observation 7c3ff951-5a41-4ef3-b9a1-8ddfa4e8a78d · outbound

This paper cites an unresolved cited work.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:17:33.911818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.357222Z digest=sha256:711c8e373ec1b1a257fef359d776b317292e8e63e26f6535d5b45de2dc3dd901

Observation d1a5bcd2-cc09-493b-a09b-067d749df2e0 · outbound

This paper cites an unresolved cited work.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:17:33.900926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.360522Z digest=sha256:ac710d77eb7761dbd34a76af84e7514f6276d699480a62e1b17f821ddc5d084e

Observation c8563fb2-e5da-44cf-bd61-a3d7a911e0e5 · outbound

This paper cites Our framework con- sists of pseudo paralinguistic label-based data condensation and LLM-based CPQA generation.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Our framework con- sists of pseudo paralinguistic label-based data condensation and LLM-based CPQA generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.364103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.364103Z digest=sha256:0715bced88145add0a83d9b11e6b0c2ef69044286f8cd9c58f0a948c58c38986

Observation 257f1e2b-b648-44d7-81eb-7239a32a713c · outbound

This paper cites an unresolved cited work.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:17:33.890556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.367822Z digest=sha256:a2e30ef93bac0b42727377db21f5016d23bb7a4a7a5137af62a34dab49e0da28

Observation a8c59dad-2984-4c64-b67b-03064de6c960 · outbound

This paper cites GPT-4 Technical Report.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation GPT-4 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.371023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.371023Z digest=sha256:881b391aded902a0a5ee28a2e14e94b8728788a269bb167a351b6d0d899ff675

Observation d4e56222-2487-41b0-b0f7-c939eed97885 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.374448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.374448Z digest=sha256:c1dcbfc552e235511a1cbbf894a886dbb644eda94551a255a72d470ddc11cb91

Observation 8ec7c9d2-eead-4631-8149-034f16ba74cb · outbound

This paper cites Qwen2-Audio Technical Report.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Qwen2-Audio Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.377815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.377815Z digest=sha256:14a2e5c363d494367e00122527592597cc7201c9e2a01de41ef11d0f5ce91569

Observation c02a3a55-2109-418b-a084-029faeb96724 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.380885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.380885Z digest=sha256:887d69b534b51e4ed71b736f17ada512074da073d887245cb96dae615ceb214d

Observation b47227b3-c8cf-413e-916b-6b0f82b780df · outbound

This paper cites MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.384225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.384225Z digest=sha256:ed563cd36f486c742401d07e96450037410eac1075ec8ecb088250ba5a7c6c57

Observation 79ba7230-e884-473a-bbec-18385fee75db · outbound

This paper cites MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.387816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.387816Z digest=sha256:cf69b3717a4d77eb6a7a2846505ac86408c960c7e2dd08314116944043b86d60

Observation c1a5cc0e-951c-4a5a-92d5-eb332a8eccf7 · outbound

This paper cites BLSP-Emo: Towards Empathetic Large Speech-Language Models.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation BLSP-Emo: Towards Empathetic Large Speech-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.391551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.391551Z digest=sha256:3925a1fee0b860f82fb2950f53007aa04a48a15bf9451b914776b4d0de49486b

Observation bcac46d7-5359-459f-a598-dae8ec7ca2c0 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.394954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.394954Z digest=sha256:d0e9abc22876f47f12c81e7db004daf421dfc91ff9424466657fa953769853c2

Observation f5e1c955-edc6-462f-bfaf-0e48780e59b4 · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.399189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.399189Z digest=sha256:b606d49b2b88c3e03e8f8b56ef8f40ce4bd66e85ee298356727361d9f1968bbe

Observation 7cd1ed08-992a-426b-a7d7-721c2ccef8fc · outbound

This paper cites Advancing large lan- guage models to capture varied speaking styles and respond prop- erly in spoken conversations,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Advancing large lan- guage models to capture varied speaking styles and respond prop- erly in spoken conversations,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.880417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.402172Z digest=sha256:7f9349b83ebc006f7291594b7fa6e1ca277a0dae5dfc5484c8f04762a586f439

Observation adf22b68-6130-4741-bcf3-9340321f1889 · outbound

This paper cites BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.405082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.405082Z digest=sha256:4a945fa36517effcfef91ffcfc97e53a69d9867287312156d043315f4b22d9ff

Observation bb7090e5-9f39-421b-b88a-f8e163372e9a · outbound

This paper cites Paralinguistics-aware speech- empowered large language models for natural conversation,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Paralinguistics-aware speech- empowered large language models for natural conversation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.868245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.409218Z digest=sha256:34447dfacece974f37b7e5a3190b47f82d6b8b282af7a3b6975a3a25ceb9d70c

Observation 6a3ad9af-4b75-44f7-8891-2fe02c825a7a · outbound

This paper cites Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.412848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.412848Z digest=sha256:4bdd14427f8b9ccc75c78545cc8f04b49a282c3d0fdf4ddf5340ccf74bbcfcd7

Observation 5cad0500-702b-4260-805d-dfb6d0be4ad9 · outbound

This paper cites AudioBench: A universal benchmark for audio large language models,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation AudioBench: A universal benchmark for audio large language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.858006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.415869Z digest=sha256:91f4f566427d446ed8276a90029a83106cd9de098c462f2fb50df2f5c17105f6

Observation 40ca13f3-6ee6-43b9-96a4-2eb4d6e4422c · outbound

This paper cites Dynamic-superb: To- wards a dynamic, collaborative, and comprehensive instruction- tuning benchmark for speech,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Dynamic-superb: To- wards a dynamic, collaborative, and comprehensive instruction- tuning benchmark for speech,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.847956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.418994Z digest=sha256:335aa60ba6a78ed9d69357083960105c911c76c3afae5ebc34351caa1c080eb9

Observation d7cda837-7c18-4c6c-822a-2c129c3e9a70 · outbound

This paper cites AIR-bench: Benchmark- ing large audio-language models via generative comprehension,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation AIR-bench: Benchmark- ing large audio-language models via generative comprehension,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.837167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.422609Z digest=sha256:e96c4cdcaebb3b81cf9139b9b60634fd179cc170e37753c36bc31efced744e3d

Observation b88e64c0-7555-4a25-9946-fbe3044dbbf5 · outbound

This paper cites Listen, think, and understand,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Listen, think, and understand,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.425641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.425641Z digest=sha256:fbc314ae9c38888d20effee28a3c9e14a6c165b6af793db5af100dbaeb706220

Observation 3c5cccfd-59b3-42fd-aa6f-9f515ba723aa · outbound

This paper cites MMAU: A mas- sive multi-task audio understanding and reasoning benchmark,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation MMAU: A mas- sive multi-task audio understanding and reasoning benchmark,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.818914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.428590Z digest=sha256:5c5ba09b3ad0ca8eb7965a9e0bdeea0ecd805204735d3c672001bbdc963ba0f0

Observation 8965e6db-6080-48ac-804c-ad5240d7ae40 · outbound

This paper cites IEMOCAP: interactive emotional dyadic motion capture database,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation IEMOCAP: interactive emotional dyadic motion capture database,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.807231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.431691Z digest=sha256:4101a8da8e72e54cd49d4ff230509f686259866935a77160a27603229c32604f

Observation d012c2ef-7d75-4efe-8248-b05d1ac94826 · outbound

This paper cites MELD: A multimodal multi-party dataset for emo- tion recognition in conversations,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation MELD: A multimodal multi-party dataset for emo- tion recognition in conversations,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.797163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.434765Z digest=sha256:5b8141c37d6200ebef3ef8fb0830a415e2221e2c1e1a4caadef8aa654f812f33

Observation a129374a-a85b-4739-941c-7dd619b5a975 · outbound

This paper cites What’s basic about basic emotions?.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation What’s basic about basic emotions?

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.785652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.438307Z digest=sha256:1dbe2f35982812cf9f7c21a400fec002f179eae6d0edc53499144da1391a0522

Observation 85b7fc7a-0eb6-4d69-b73d-087401a787b0 · outbound

This paper cites Theories of emotion,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Theories of emotion,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.775706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.441242Z digest=sha256:521786acc0b60ae3d27abb09c41995838b5c4381b5c2ea84e7c304aa3f375970

Observation 2418c4d2-380d-4b51-8d7d-6902ebc7b32e · outbound

This paper cites EmoBox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation EmoBox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.765848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.444355Z digest=sha256:143656e4b719e354f2d4c09dd54d162ccfdc7a09743fb46cb89183f225f7b94c

Observation 8bc8255e-2381-4b4b-b1f3-04acb1b2d49b · outbound

This paper cites Evidence for a three-factor the- ory of emotions,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Evidence for a three-factor the- ory of emotions,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.755021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.447515Z digest=sha256:f8bb1495b67a048f096214ba8249ed08a3db5c3d920ef29220b1e4cf79135014

Observation 5fc62db6-16d6-42cf-accd-f389cbbbabf2 · outbound

This paper cites Goemotions: A dataset of fine-grained emo- tions,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Goemotions: A dataset of fine-grained emo- tions,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.745284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.450756Z digest=sha256:2223ab9c3f605e6a4a23ef33e56f0ec158bfc25c49d1956e31229d0cfbd99318

Observation 90ead142-5439-4f7a-bcd5-0cfd198265af · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation emotion2vec: Self-supervised pre-training for speech emotion representation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.734700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.454003Z digest=sha256:28e9d8ef0a10cd9821af50826b242cfe54d1de32ff328c3bd8bb6cdc9259daae

Observation 5d0fc0a3-4a9a-4c29-836a-9cf4ef2e8d60 · outbound

This paper cites Dawn of the trans- former era in speech emotion recognition: Closing the valence gap,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Dawn of the trans- former era in speech emotion recognition: Closing the valence gap,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.457390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.457390Z digest=sha256:bfc39874fafdcd65a3873f45e7e465c1beaec6081a3a4f0311735778ac74af9a

Observation 481c58b7-2554-485f-b02e-aeb833128bff · outbound

This paper cites Building naturalistic emotionally bal- anced speech corpus by retrievingemotional speech from existing podcast recordings,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Building naturalistic emotionally bal- anced speech corpus by retrievingemotional speech from existing podcast recordings,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.718410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.460553Z digest=sha256:a6c07fdbc1cad1a2b4fb701769f1d10c4cb3befb8aef9e67ac6b14e6ddf9e935

Observation 3cbdafd9-b7a6-4a4b-ac0c-3a9ef94c2ea6 · outbound

This paper cites WavLM: Large-scale self- supervised pre-training for full stack speech processing,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation WavLM: Large-scale self- supervised pre-training for full stack speech processing,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.463514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.463514Z digest=sha256:d93d8a6682b6b258d3ac689c56b2f799d73bb355eaf8142554e194d3a53f327c

Observation 26111fcc-14e8-4c9d-8386-3f5270689267 · outbound

This paper cites Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.701994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.466518Z digest=sha256:ad3f3f296de52f5d4493593725010659b77125f8f741316d4a031b1f54e23584

Observation 35d844f6-a36d-41b0-a4a0-63ae8a49271c · outbound

This paper cites V oxCeleb2: Deep speaker recognition,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation V oxCeleb2: Deep speaker recognition,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:33.469808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:33.469808Z digest=sha256:d7134d41f5ec7fb0eccc6a252fa775c84b4cd61060be683a5714f1bbfcc013a7

Observation 3a6e8142-01c7-4b17-8889-90906e41acac · outbound

This paper cites WhisperX: Time- accurate speech transcription of long-form audio,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation WhisperX: Time- accurate speech transcription of long-form audio,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.686760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.472777Z digest=sha256:1c372e0aa1a242b10d4773c27b2530449deb056abf6ca1c7f431b1f40c2a22a3

Observation e96a01a8-e897-4f96-8521-8b45b64ba23a · outbound

This paper cites The Llama 3 herd of models,.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation The Llama 3 herd of models,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:17:33.676651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.475816Z digest=sha256:62950bc122509502fbe422a4c8fea8a76712aba5affbc234e008287ff3ba1e61

Pith citing papers

Observation 5b3e0b13-e110-4c94-b1d6-097889e13798 · inbound

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation cites this paper.

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:17:33.666296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:17:33.349929Z digest=sha256:c81d416c1e984852fd97c7b0133f052b89fdd88a2f1df8078e57c47e69d99767