Pith. sign in

Paper Citation Record · LEDGER

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval

As of 20 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2505.19437.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19437 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:17:56.416700Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:17:54.000523Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:17:56.702242Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a4d2841a-34d1-4abc-8a00-52bae8c84bb5 · outbound

This paper cites RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:17:56.775310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:17:54.000523Z digest=sha256:c89fc428fff8855c0fe36ffd7763bf33755bfaa096fa497d3c16c967e8293cbb

Observation 4fd6735a-f6ad-48e7-9a57-56f91ff8816f · outbound

This paper cites 1, the proposed RA-CLAP model is designed to learn a joint representation of speech and text through a con- trastive learning and self-distillation framework.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval 1, the proposed RA-CLAP model is designed to learn a joint representation of speech and text through a con- trastive learning and self-distillation framework

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:01.073427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:17:54.068751Z digest=sha256:0cc1b9afcabe100c11e28f77d52b4afd287dbeeb0c8e062cbe58318960e99806

Observation 255fd815-9fcf-4735-8fdd-12a953119b6c · outbound

This paper cites an unresolved cited work.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:18:00.833157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:17:54.170786Z digest=sha256:2f2986d5b07aa6e7e39ff08163ceeaa64573e54ea84bff4015eff2b0f27bdd82

Observation 22cfcac5-4ecd-4f32-a2f6-cd10482d2d70 · outbound

This paper cites an unresolved cited work.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:18:00.385503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:17:54.269768Z digest=sha256:084e987bbd55a196ee99118860ee05ca08d88f1ecdd68e547e777b363e7b33c3

Observation dcab4470-541c-4159-ab6e-9a460c34c306 · outbound

This paper cites Multi-level knowledge distillation for speech emotion recogni- tion in noisy conditions,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Multi-level knowledge distillation for speech emotion recogni- tion in noisy conditions,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:00.162738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:17:54.341870Z digest=sha256:b93ed102bd7e78da1325d5a9927939514238627018d932f851fbe8cc28142584

Observation e2d36bc6-127e-4640-b312-de50206a9f3c · outbound

This paper cites Iterative prototype refinement for ambiguous speech emotion recognition,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Iterative prototype refinement for ambiguous speech emotion recognition,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:59.864218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:17:54.433356Z digest=sha256:a68349af70c5f2cd45700022e52b54501083593ab02c41e63d260216749b3207

Observation 8dcd976f-92b6-47d7-ba03-c5630c96c011 · outbound

This paper cites Fine- grained disentangled representation learning for multimodal emo- tion recognition,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Fine- grained disentangled representation learning for multimodal emo- tion recognition,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:59.534203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:17:54.525076Z digest=sha256:1ddb5265dd959685d033fdfdfe6923f8abc3ce8a9e557566497082482fc1a9b1

Observation 7b381f07-e144-40b8-b0d9-714a6aeccfdf · outbound

This paper cites Enhancing multimodal emotion recognition through multi- granularity cross-modal alignment,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Enhancing multimodal emotion recognition through multi- granularity cross-modal alignment,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:59.287679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:17:54.622654Z digest=sha256:75b2fdaa3767ff323218cfb3efa86e3cbdcbadf4d26326484d6d2039efc3e4d6

Observation 24a5fce8-6dbb-496e-a982-5e77b7170dd9 · outbound

This paper cites Enhancing emotion recognition in incomplete data: A novel cross-modal alignment, reconstruction, and refinement framework,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Enhancing emotion recognition in incomplete data: A novel cross-modal alignment, reconstruction, and refinement framework,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:58.971264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:17:54.695780Z digest=sha256:31ade359a99cc8c3dc07f90820f583a47f9f036bfbe83c783bda9f7bd387f707

Observation b5e1bef2-6569-419b-a241-a2292c9f7476 · outbound

This paper cites Emotion- preserving prosody anonymization network for voice privacy pro- tection,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Emotion- preserving prosody anonymization network for voice privacy pro- tection,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:58.689480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:17:54.769076Z digest=sha256:45987ad2e071bb737496cce469c7bd562bb6b17389a68e2d13ca672696124a81

Observation 7b76d34f-d7e0-4a80-be68-e6aa57a9c381 · outbound

This paper cites Speaker-Text Retrieval via Contrastive Learning.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Speaker-Text Retrieval via Contrastive Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:54.860761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:54.860761Z digest=sha256:857e2eeac22b10a0a410d0d587d50a10f509d138a8cf280a31019a1c5ce245b8

Observation 5c13774c-0ad4-4b6f-be57-d16cd86060c4 · outbound

This paper cites ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:54.955489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:54.955489Z digest=sha256:3cab2cd49e1284b54513174b8fd87db0a6cb1620798058d9e6d78314e6a1fd34

Observation 885dc978-f9c3-4645-afed-bb1808695639 · outbound

This paper cites Prompttts: Control- lable text-to-speech with text descriptions,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Prompttts: Control- lable text-to-speech with text descriptions,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.047419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.047419Z digest=sha256:fc9d02b9b8a816ffdd95c48e35e56511defb3bf47f0cda69a1f7e691c48cc952

Observation 2dde2d89-3f11-4d8f-8a91-71334313d05a · outbound

This paper cites Prompttts++: Controlling speaker identity in prompt-based text-to-speech using natural language descriptions,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Prompttts++: Controlling speaker identity in prompt-based text-to-speech using natural language descriptions,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.119979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.119979Z digest=sha256:c3861b75e63aec5750de0799aaeffe38ef9f2ee93ba8977ad5458be0624baa70

Observation 217c859d-ace3-4f9b-935d-7b0465d73791 · outbound

This paper cites Stylecap: Automatic speaking-style captioning from speech based on speech and lan- guage self-supervised learning models,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Stylecap: Automatic speaking-style captioning from speech based on speech and lan- guage self-supervised learning models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:58.439368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:17:55.198369Z digest=sha256:130e2209e7aca38ffdbf198d5aeb7f7bb35cdea8897937a585a7bd07da5c6baf

Observation eecc696b-4dec-42a6-86ce-e67a873841e5 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Learning transferable visual models from natural language supervision,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.282718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.282718Z digest=sha256:f7e9e1e429718343b0f75d80ec1346c4efab7d3829b15742a766877b852ce3cd

Observation 58ce6195-63f3-4120-a029-d8cfed917ff5 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.352417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.352417Z digest=sha256:95c6053cb93badb308b24b717557b7d7d2ec8edc63745dbbd9858b5f80c41b55

Observation 1ec33af7-db3f-4148-acc8-169c6d83f315 · outbound

This paper cites Gemo-clap: Gender-attribute-enhanced contrastive language- audio pretraining for accurate speech emotion recognition,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Gemo-clap: Gender-attribute-enhanced contrastive language- audio pretraining for accurate speech emotion recognition,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.424565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.424565Z digest=sha256:c647550c3dcfad5b66f985c8dd7ed25fb64d3b33d7c12d851d6eb1fcda8aec24

Observation 0bf44369-fa18-4cc9-92a8-1c3955fd0b31 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.489986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.489986Z digest=sha256:031d8bdd274638c2d22c5e338c334a4420315bd8587ad177d92af98d0d9f77af

Observation b98c4c3d-eb89-4736-9d6d-234d00ccfc05 · outbound

This paper cites Textrolspeech: A text style control speech corpus with codec language text-to-speech models,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Textrolspeech: A text style control speech corpus with codec language text-to-speech models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.585634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.585634Z digest=sha256:9185826b5326a801c94cf4aa5adf6667b3c2b8c09f9f8c34a588331dde934437

Observation 03232a1d-a6c3-40eb-88ec-7f590001e8c1 · outbound

This paper cites Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (ver- sion 0.92),.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (ver- sion 0.92),

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.683776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.683776Z digest=sha256:fe6604358df9b01865dbc1f5d1e7ae299678dad205d9add2cad6a8eb4101978e

Observation b41c8ce9-08fa-448e-9a36-22230ca671ba · outbound

This paper cites Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:58.010603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:17:55.766272Z digest=sha256:13841c5a87983df0100c7bbe9a85b4fcf7290426ad57f89a4d1a18a3b665e915

Observation bc78f9b0-6e12-40ea-8dd1-b6dfefc3985e · outbound

This paper cites Toronto emotional speech set (tess)-younger talker happy,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Toronto emotional speech set (tess)-younger talker happy,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.841134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.841134Z digest=sha256:cf674b76622e0be0479f91e1a2ee796cc7cc35531a490d419e38154340ee64eb

Observation 926bbc5f-94fc-49f7-98c5-5813ee848590 · outbound

This paper cites Mead: A large-scale audio-visual dataset for emotional talking-face generation,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Mead: A large-scale audio-visual dataset for emotional talking-face generation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:55.935541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:55.935541Z digest=sha256:99ec0168275e1d3078a743a982f05a336c1aea126d082520eb57b30ddcc04f19

Observation f03dea00-5445-4c91-a680-281a4ca24cea · outbound

This paper cites Speaker-dependent audio- visual emotion recognition.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Speaker-dependent audio- visual emotion recognition

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:57.693973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:17:56.006922Z digest=sha256:17b1be1693f0d742a433a5521da873d203a8c7c6c72a8acfc3d4791896cfe6ab

Observation 4d642703-dd2c-4e55-aa53-48d00d88f84c · outbound

This paper cites Categorical and dimensional ratings of emotional speech: Behavioral findings from the morgan emotional speech set,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Categorical and dimensional ratings of emotional speech: Behavioral findings from the morgan emotional speech set,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:57.338662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:17:56.074876Z digest=sha256:812e1c81dfe9e14f6a65519d379eea334bd2f275c47522095eae12f16b932f34

Observation 3f5edecf-56f6-4bc2-b46c-2e855d6fd05c · outbound

This paper cites Speechcraft: A fine-grained expressive speech dataset with natural language description,.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval Speechcraft: A fine-grained expressive speech dataset with natural language description,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:17:57.013435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:17:56.143003Z digest=sha256:7a0df16edb917990bac00d1a6a16e5957ecd7503c57def3ddb77d767fa00be93

Observation 72b3b378-ef9c-4449-8708-5f8c65c46d46 · outbound

This paper cites AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:56.239809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:56.239809Z digest=sha256:e0e5df3a4aa49a2c60c8f1287f6ec918b641764153b8aa14dc81659ca39a7166

Observation b6c69078-b779-436a-8e7e-6dc2e647e0ad · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:56.327895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:56.327895Z digest=sha256:1f810b9a2708864ad75228f5fbdcc3be83ab7f294d0332586cc745553854161a

Observation fea97d2f-3f14-46fd-8cc3-9a25b645b90f · outbound

This paper cites LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:56.416700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:56.416700Z digest=sha256:273cec6119a431e042beddbfa3ad167e55f03bd332dcd89d9497e656fcffb9f7

Pith citing papers

Observation a4d2841a-34d1-4abc-8a00-52bae8c84bb5 · inbound

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval cites this paper.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:17:56.775310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:17:54.000523Z digest=sha256:c89fc428fff8855c0fe36ffd7763bf33755bfaa096fa497d3c16c967e8293cbb