Pith. sign in

REVIEW 1 cited by

DefSent: Sentence Embeddings using Definition Sentences

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.04339 v3 pith:HJTKQ5RV submitted 2021-05-10 cs.CL

classification cs.CL
keywords datasetsdefsentmethodstasksavailablesentencebettercomparably
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sentence embedding methods using natural language inference (NLI) datasets have been successfully applied to various tasks. However, these methods are only available for limited languages due to relying heavily on the large NLI datasets. In this paper, we propose DefSent, a sentence embedding method that uses definition sentences from a word dictionary, which performs comparably on unsupervised semantics textual similarity (STS) tasks and slightly better on SentEval tasks than conventional methods. Since dictionaries are available for many languages, DefSent is more broadly applicable than methods using NLI datasets without constructing additional datasets. We demonstrate that DefSent performs comparably on unsupervised semantics textual similarity (STS) tasks and slightly better on SentEval tasks to the methods using large NLI datasets. Our code is publicly available at https://github.com/hpprc/defsent .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval

    cs.IR 2025-06 conditional novelty 6.0 of 10

    Vela adapts an audio MLLM into a universal text-audio embedding model using 'in one word' prompts, in-context examples, and text-only contrastive training, outperforming CLAP-style models on retrieval benchmarks.

Pith tools