Pith. sign in

REVIEW 1 cited by

Self-supervised Contrastive Cross-Modality Representation Learning for Spoken Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.03381 v1 pith:KLZ6FPLP submitted 2021-09-08 cs.CL cs.AIcs.LGcs.SDeess.AS

classification cs.CLcs.AIcs.LGcs.SDeess.AS
keywords questionself-supervisedspokenansweringcontrastivemodelproposestage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Spoken question answering (SQA) requires fine-grained understanding of both spoken documents and questions for the optimal answer prediction. In this paper, we propose novel training schemes for spoken question answering with a self-supervised training stage and a contrastive representation learning stage. In the self-supervised stage, we propose three auxiliary self-supervised tasks, including utterance restoration, utterance insertion, and question discrimination, and jointly train the model to capture consistency and coherence among speech documents without any additional data or annotations. We then propose to learn noise-invariant utterance representations in a contrastive objective by adopting multiple augmentation strategies, including span deletion and span substitution. Besides, we design a Temporal-Alignment attention to semantically align the speech-text clues in the learned common space and benefit the SQA tasks. By this means, the training schemes can more effectively guide the generation model to predict more proper answers. Experimental results show that our model achieves state-of-the-art results on three SQA benchmarks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HNCSE: Advancing Sentence Embeddings via Hybrid Contrastive Learning with Hard Negatives

    cs.CL 2024-11 reject novelty 4.0 of 10

    HNCSE reports 2-point average STS gains over SimCSE using positive mixing and hard-negative mixing, but the method is under-specified and unverified.

Pith tools