OmniScore is a family of lightweight deterministic learned metrics that approximate LLM-judge behavior for reliable multilingual evaluation of generative text in tasks such as QA, translation, and summarization.
Time to impeach LLM -as-a-judge: Programs are the future of evaluation
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CL 2years
2026 2representative citing papers
Judge language and model backbone interact so strongly that backbone rankings invert across languages, and no single judge backbone wins in all five languages tested.
citing papers explorer
-
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
OmniScore is a family of lightweight deterministic learned metrics that approximate LLM-judge behavior for reliable multilingual evaluation of generative text in tasks such as QA, translation, and summarization.
-
Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation
Judge language and model backbone interact so strongly that backbone rankings invert across languages, and no single judge backbone wins in all five languages tested.