Pith. sign in

REVIEW 4 cited by

Sabi\'a-3 Technical Report

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.12049 v4 pith:BAJ7IIXD submitted 2024-10-15 cs.CL cs.AI

classification cs.CLcs.AI
keywords sabilargemodelperformancereporttasksacademicacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This report presents Sabi\'a-3, our new flagship language model, and Sabiazinho-3, a more cost-effective sibling. The models were trained on a large brazilian-centric corpus. Evaluations across diverse professional and academic benchmarks show a strong performance on Portuguese and Brazil-related tasks. Sabi\'a-3 shows large improvements in comparison to our previous best of model, Sabia-2 Medium, especially in reasoning-intensive tasks. Notably, Sabi\'a-3's average performance matches frontier LLMs, while it is offered at a three to four times lower cost per token, reinforcing the benefits of domain specialization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A new benchmark, TiEBe, measures LLM recall of notable events across time, regions, and languages, and finds large geographic disparities correlated with GDP, HDI, and schooling.

  2. Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Study

    cs.CY 2025-12 unverdicted novelty 5.0 of 10

    In interviews with 11 Portuguese-language model developers, four AI ethics tools guided general ethical reflection but failed to surface Portuguese-specific harms like cultural misrepresentation and low language performance.

  3. BRoverbs -- Measuring how much LLMs understand Portuguese proverbs

    cs.CL 2025-09 conditional novelty 5.0 of 10

    BRoverbs lets researchers test whether language models understand Portuguese proverbs; commercial models nearly master it, small models often guess randomly.

  4. BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning

    cs.CL 2025-08 conditional novelty 4.0 of 10

    An updated BLUEX benchmark with 1,422 questions and GPT-4o-generated captions that make image-based questions usable by text-only LLMs.

Pith tools