Pith. sign in

Paper Citation Record · LEDGER

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval

As of 7 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2505.19437.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19437 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:17:56.416700Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:17:54.000523Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:17:56.702242Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a4d2841a-34d1-4abc-8a00-52bae8c84bb5 · outbound

This paper cites RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:17:56.775310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:17:54.000523Z digest=sha256:e39b39efbff6952e672f1cfe6944d2c7d064266de2c4f3beb1c293182bb5380a

Observation 4fd6735a-f6ad-48e7-9a57-56f91ff8816f · outbound

This paper cites 1, the proposed RA-CLAP model is designed to learn a joint representation of speech and text through a con- trastive learning and self-distillation framework.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval 1, the proposed RA-CLAP model is designed to learn a joint representation of speech and text through a con- trastive learning and self-distillation framework

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:01.073427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:17:54.068751Z digest=sha256:1521b2ccf53536f5fc31f5e823bd7b07c9868be6bd5cc20ab12516158d2d4497

Observation 255fd815-9fcf-4735-8fdd-12a953119b6c · outbound

This paper cites an unresolved cited work.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:18:00.833157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:17:54.170786Z digest=sha256:a0320885b96ee9a90ebf1f9b2a566e0a9318c80fe1717fea0fd667579c469790

Observation 22cfcac5-4ecd-4f32-a2f6-cd10482d2d70 · outbound

This paper cites an unresolved cited work.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:18:00.385503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:17:54.269768Z digest=sha256:9caf4d94e7c9681404fd4231359ac74d2caece79a6ea967f1b6d42bfd682d4b9

Observation dcab4470-541c-4159-ab6e-9a460c34c306 · outbound

This paper cites Multi-level knowledge distillation for speech emotion recogni- tion in noisy conditions,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Multi-level knowledge distillation for speech emotion recogni- tion in noisy conditions,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:00.162738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:17:54.341870Z digest=sha256:d196f63a4fc70e32b06ae60dbafa64c4ee14ff6e4d915fc845f7f2c2d00def8d

Observation e2d36bc6-127e-4640-b312-de50206a9f3c · outbound

This paper cites Iterative prototype refinement for ambiguous speech emotion recognition,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Iterative prototype refinement for ambiguous speech emotion recognition,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:59.864218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:17:54.433356Z digest=sha256:dd9422fda0c7a1bcec9e16d7b3022fefd4fc9cbfe622619fb85462f43b37bba5

Observation 8dcd976f-92b6-47d7-ba03-c5630c96c011 · outbound

This paper cites Fine- grained disentangled representation learning for multimodal emo- tion recognition,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Fine- grained disentangled representation learning for multimodal emo- tion recognition,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:59.534203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:17:54.525076Z digest=sha256:f928f151ac0a46cafb15c44282ae63556ff12be053bc899f0d4fc5cc8b1cafae

Observation 7b381f07-e144-40b8-b0d9-714a6aeccfdf · outbound

This paper cites Enhancing multimodal emotion recognition through multi- granularity cross-modal alignment,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Enhancing multimodal emotion recognition through multi- granularity cross-modal alignment,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:59.287679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:17:54.622654Z digest=sha256:ac3a2076f681c072b815bc5588f4e8dc28528b8a840cc8551f56ecc42b4c3fd6

Observation 24a5fce8-6dbb-496e-a982-5e77b7170dd9 · outbound

This paper cites Enhancing emotion recognition in incomplete data: A novel cross-modal alignment, reconstruction, and refinement framework,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Enhancing emotion recognition in incomplete data: A novel cross-modal alignment, reconstruction, and refinement framework,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:58.971264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:17:54.695780Z digest=sha256:40dda133d566835706829430c400c243380074b069fdc659aecf18f26fb6b4a7

Observation b5e1bef2-6569-419b-a241-a2292c9f7476 · outbound

This paper cites Emotion- preserving prosody anonymization network for voice privacy pro- tection,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Emotion- preserving prosody anonymization network for voice privacy pro- tection,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:58.689480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:17:54.769076Z digest=sha256:8115ad7b1fab87db6bd13ca7315550a0380771f218ecbc0ae9d08e7feed5baae

Observation 7b76d34f-d7e0-4a80-be68-e6aa57a9c381 · outbound

This paper cites Speaker-Text Retrieval via Contrastive Learning.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Speaker-Text Retrieval via Contrastive Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:54.860761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:54.860761Z digest=sha256:5e2d0536ca5852e06e6e5c14f5cdc1c945e93e9d61713c4e3cc2cedb9bb0e637

Observation 5c13774c-0ad4-4b6f-be57-d16cd86060c4 · outbound

This paper cites ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:54.955489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:54.955489Z digest=sha256:c662b11828f61d3f82ca5c7818f45fd5816bb2c67da3c5e3b95284d17fe84fc8

Observation 885dc978-f9c3-4645-afed-bb1808695639 · outbound

This paper cites Prompttts: Control- lable text-to-speech with text descriptions,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Prompttts: Control- lable text-to-speech with text descriptions,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.047419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.047419Z digest=sha256:5b55de5f04bcfc0439776b94faf788b38cc06dcfcd9a33f184725f15b0863831

Observation 2dde2d89-3f11-4d8f-8a91-71334313d05a · outbound

This paper cites Prompttts++: Controlling speaker identity in prompt-based text-to-speech using natural language descriptions,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Prompttts++: Controlling speaker identity in prompt-based text-to-speech using natural language descriptions,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.119979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.119979Z digest=sha256:91a7ca24b4360b35be24ca8c3e5b2c88d192194ac1a433739cecbb9a6f453f52

Observation 217c859d-ace3-4f9b-935d-7b0465d73791 · outbound

This paper cites Stylecap: Automatic speaking-style captioning from speech based on speech and lan- guage self-supervised learning models,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Stylecap: Automatic speaking-style captioning from speech based on speech and lan- guage self-supervised learning models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:58.439368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:17:55.198369Z digest=sha256:885d390aff6128f986e7169183559389679933e9fb1d2ec554160f96d8467991

Observation eecc696b-4dec-42a6-86ce-e67a873841e5 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Learning transferable visual models from natural language supervision,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.282718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.282718Z digest=sha256:6fec5a2d0059a6b916c7ab8cbf33a46249bc7c2d01991ca1de573ecba64f924c

Observation 58ce6195-63f3-4120-a029-d8cfed917ff5 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.352417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.352417Z digest=sha256:ba8a97072202cb78965636cd9a1e91caa837680b3b3599add05a8cd4d43608e6

Observation 1ec33af7-db3f-4148-acc8-169c6d83f315 · outbound

This paper cites Gemo-clap: Gender-attribute-enhanced contrastive language- audio pretraining for accurate speech emotion recognition,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Gemo-clap: Gender-attribute-enhanced contrastive language- audio pretraining for accurate speech emotion recognition,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.424565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.424565Z digest=sha256:4b71cef8d1e569d57bc5a3bb31c151a18ce2588c04f7c4c9525d899689f5e0b7

Observation 0bf44369-fa18-4cc9-92a8-1c3955fd0b31 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.489986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.489986Z digest=sha256:91c677fa418e616c00358ce395701b893482356d0e86cad5e2b9cc4ab17c9b68

Observation b98c4c3d-eb89-4736-9d6d-234d00ccfc05 · outbound

This paper cites Textrolspeech: A text style control speech corpus with codec language text-to-speech models,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Textrolspeech: A text style control speech corpus with codec language text-to-speech models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.585634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.585634Z digest=sha256:c097475a626983f292bcaf7ef49d7e867e7a5e153855498cbd67fb39b48a6b62

Observation 03232a1d-a6c3-40eb-88ec-7f590001e8c1 · outbound

This paper cites Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (ver- sion 0.92),.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (ver- sion 0.92),

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.683776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.683776Z digest=sha256:e3eb1325f5f6c52aa7d813751b95c69b65f740615366b2217283c6f95854d22e

Observation b41c8ce9-08fa-448e-9a36-22230ca671ba · outbound

This paper cites Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:58.010603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:17:55.766272Z digest=sha256:322559688433e152141a5b61849e4d01e1baa9e79c5272d0486fef3770c126fa

Observation bc78f9b0-6e12-40ea-8dd1-b6dfefc3985e · outbound

This paper cites Toronto emotional speech set (tess)-younger talker happy,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Toronto emotional speech set (tess)-younger talker happy,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.841134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.841134Z digest=sha256:24163e5192436250f17eec60f980886d70cc7bc18b6e7b7552f1588e0177972f

Observation 926bbc5f-94fc-49f7-98c5-5813ee848590 · outbound

This paper cites Mead: A large-scale audio-visual dataset for emotional talking-face generation,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Mead: A large-scale audio-visual dataset for emotional talking-face generation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.935541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.935541Z digest=sha256:31f6eb88507922c7fe5ee24fa57f71a727d4346684252eebc54192ce035fef59

Observation f03dea00-5445-4c91-a680-281a4ca24cea · outbound

This paper cites Speaker-dependent audio- visual emotion recognition.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Speaker-dependent audio- visual emotion recognition

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:57.693973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:17:56.006922Z digest=sha256:b9a7b6cedd7a32b2020978af6fb91c05f3a2983d37b11167b6a53a163bcfc123

Observation 4d642703-dd2c-4e55-aa53-48d00d88f84c · outbound

This paper cites Categorical and dimensional ratings of emotional speech: Behavioral findings from the morgan emotional speech set,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Categorical and dimensional ratings of emotional speech: Behavioral findings from the morgan emotional speech set,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:57.338662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:17:56.074876Z digest=sha256:e5c081b5eceeb8ab9e85c4631adba0a0d5add98a25a7d2d0ebb638e993d62aba

Observation 3f5edecf-56f6-4bc2-b46c-2e855d6fd05c · outbound

This paper cites Speechcraft: A fine-grained expressive speech dataset with natural language description,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Speechcraft: A fine-grained expressive speech dataset with natural language description,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:57.013435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:17:56.143003Z digest=sha256:2e7f0b4734f47afb0517bd1e94cce3180ab45ab33431d2ef39d0dbc7f507d083

Observation 72b3b378-ef9c-4449-8708-5f8c65c46d46 · outbound

This paper cites AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:56.239809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:56.239809Z digest=sha256:06024bcd47e4a2018f10f785e49825e46bcbf7a2428b6726b12f948bd4acb0ea

Observation b6c69078-b779-436a-8e7e-6dc2e647e0ad · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:56.327895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:56.327895Z digest=sha256:eb7e72a96da73dcca0ccd0a3bb91b9a91801bc5b3cac5dc1fa83a5f03b0defc9

Observation fea97d2f-3f14-46fd-8cc3-9a25b645b90f · outbound

This paper cites LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:56.416700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:56.416700Z digest=sha256:23b1b9a6097cbfe5e564d3e484c02c296c017b42965560950a5a2a79a267fffd

Pith citing papers

Observation a4d2841a-34d1-4abc-8a00-52bae8c84bb5 · inbound

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval cites this paper.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:17:56.775310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:17:54.000523Z digest=sha256:e39b39efbff6952e672f1cfe6944d2c7d064266de2c4f3beb1c293182bb5380a