Pith. sign in

REVIEW 2 cited by

Sub-Sentence Encoder: Contrastive Learning of Propositional Semantic Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.04335 v1 pith:ORNHLZFT submitted 2023-11-07 cs.CL cs.AI

classification cs.CLcs.AI
keywords sub-sentencetextsemanticembeddingsencoderencodersatomiccontextual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce sub-sentence encoder, a contrastively-learned contextual embedding model for fine-grained semantic representation of text. In contrast to the standard practice with sentence embeddings, where the meaning of an entire sequence of text is encoded into a fixed-length vector, the sub-sentence encoder learns to produce distinct contextual embeddings corresponding to different atomic propositions, i.e. atomic units of meaning expressed within a text sequence. The sub-sentence embeddings are contrastively learned to recognize (inferred) semantic equivalence between propositions across different text sequences. Our experiments show the effectiveness of sub-sentence encoders in applications, such as retrieving supporting facts for fine-grained text attribution or recognizing the conditional semantic similarity between texts. In practice, we demonstrate that sub-sentence encoders keep the same level of inference cost and space complexity compared to sentence encoders.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM

    cs.CL 2024-12 reject novelty 5.0 of 10

    SumAutoEval is an LLM-based entity-level summarization evaluator with four dimensions; its claimed human-correlation advantage is not consistently supported by the experiments.

  2. DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Factuality scores for long-form text depend heavily on how decomposition and decontextualization are ordered; the new DnDScore verifies atomic subclaims together with their decontextualized context.

Pith tools