Pith. sign in

REVIEW 4 cited by

SwiLTra-Bench: The Swiss Legal Translation Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.01372 v2 pith:GKAHH2O4 submitted 2025-03-03 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords translationlegalswissacrossbenchmarkbestevaluationexpert
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In Switzerland legal translation is uniquely important due to the country's four official languages and requirements for multilingual legal documentation. However, this process traditionally relies on professionals who must be both legal experts and skilled translators -- creating bottlenecks and impacting effective access to justice. To address this challenge, we introduce SwiLTra-Bench, a comprehensive multilingual benchmark of over 180K aligned Swiss legal translation pairs comprising laws, headnotes, and press releases across all Swiss languages along with English, designed to evaluate LLM-based translation systems. Our systematic evaluation reveals that frontier models achieve superior translation performance across all document types, while specialized translation systems excel specifically in laws but under-perform in headnotes. Through rigorous testing and human expert validation, we demonstrate that while fine-tuning open SLMs significantly improves their translation quality, they still lag behind the best zero-shot prompted frontier models such as Claude-3.5-Sonnet. Additionally, we present SwiLTra-Judge, a specialized LLM evaluation system that aligns best with human expert assessments.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries

    cs.CY 2026-08 conditional novelty 6.0 of 10

    Out-of-the-box LLMs can match or beat top human candidates on Italian bar and judge essay exams but all fail the notary exam, which requires constrained legal drafting and planning.

  2. The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Reasoning traces improve legal MT mainly when enabled at inference, and training with reasoning keeps those traces compact enough to be cost-effective.

  3. TransLaw: A Large-Scale Dataset and Multi-Agent Benchmark Simulating Professional Translation of Hong Kong Case Law

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A seven-agent LLM translation framework and a 344-judgment English-to-Chinese dataset outperform single-agent baselines on legal accuracy, but not on human style.

  4. Reasoning Before Translation: Enhancing Legal Machine Translation with Structured Reasoning

    cs.CL 2026-07 conditional novelty 5.0 of 10

    On Swiss legal translation, reinforcement learning with a ChrF reward improves small open models more than supervised fine-tuning, but frontier reasoning models still score higher.

Pith tools