REVIEW 1 cited by
Open Sentence Embeddings for Portuguese with the Serafim PT* encoders family
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Sentence encoder encode the semantics of their input, enabling key downstream applications such as classification, clustering, or retrieval. In this paper, we present Serafim PT*, a family of open-source sentence encoders for Portuguese with various sizes, suited to different hardware/compute budgets. Each model exhibits state-of-the-art performance and is made openly available under a permissive license, allowing its use for both commercial and research purposes. Besides the sentence encoders, this paper contributes a systematic study and lessons learned concerning the selection criteria of learning objectives and parameters that support top-performing encoders.
Forward citations
Cited by 1 Pith paper
-
MTEB-BR: A Text Embedding Benchmark for Brazilian Portuguese
A native 22-task Brazilian-Portuguese embedding benchmark cleanly tiers 93 models, places an open model in the unresolved top tier, and finds only moderate rank correlation (ρ=0.75) with the multilingual MTEB board.
Discussion (0). Continue with ORCID to comment.