Pith. sign in

Paper Citation Record · LEDGER

MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2410.17578.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.17578 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:28:19.967202Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7a29ac10-0039-49b7-944c-d671fd8dc7a5 · inbound

A 2-step Framework for Automated Literary Translation Evaluation: Its Promises and Pitfalls cites this paper.

A 2-step Framework for Automated Literary Translation Evaluation: Its Promises and Pitfalls MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T04:31:46.372508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:31:46.372508Z digest=sha256:177b77d9b945a13873c8d2f911616ba142bf8fc5940e0d1fc37079f5d9106b05

Observation aaaaff09-b013-4962-823d-53722368e49a · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models

Reference 209

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:34.764278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:2e252648c54e9e9b9d6d6e9c3f808d60e320d795ad7e69acd67f3321137aa82d

Observation 0cbd0ccb-bf5e-4f99-873c-1821f29e00b6 · inbound

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model cites this paper.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.878653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.878653Z digest=sha256:ce30a8bcbd032dfe3efab5185f3d7fec04a68190bec694771e8a629077f10c9e

Observation 0408c3b9-bd5e-462c-baf9-d0e017680522 · inbound

ExTrans: Multilingual Deep Reasoning Translation via Exemplar-Enhanced Reinforcement Learning cites this paper.

ExTrans: Multilingual Deep Reasoning Translation via Exemplar-Enhanced Reinforcement Learning MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:28:19.967202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:28:19.967202Z digest=sha256:8228ba71ddd239069c8084ff21aff3914665c9dc13b0f66a3b8b9256619a492f

Observation a5a099c4-8d9f-457a-9cf1-4728329dce4f · inbound

Controlling Language Confusion in Multilingual LLMs cites this paper.

Controlling Language Confusion in Multilingual LLMs MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:09.535726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:26:09.535726Z digest=sha256:5fa3a683b5534a3f4c8b0df80a56250e93bad262611bb5f984e7260949fe704c

Observation 6f53474a-1f8d-47e8-b372-0baa0c5980b3 · inbound

IMPACT: Inflectional Morphology Probes Across Complex Typologies cites this paper.

IMPACT: Inflectional Morphology Probes Across Complex Typologies MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:13.985961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:13.985961Z digest=sha256:bb4c6fd715384c7d51d577833854aac53ca8246067d023416371706d8db9b976

Observation 80d65daf-bf30-4c19-a3bb-6c5f5417cd2a · inbound

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models cites this paper.

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:55.894317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:55.894317Z digest=sha256:9792e91e2289764faf369fa81f9285e095b6a2d05fc78fa27d5189421d262fab

Observation 55aa138a-6109-4068-afdf-a06fcc171e31 · inbound

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation cites this paper.

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:05:45.566099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T20:02:40.980538Z digest=sha256:ff024cf9cc247abde81ce4e3612c9f14e85b7a56d2d301199782d4bca55d5edb

Observation 1ac2f856-344c-4bd4-b0fc-19e7e828ad40 · inbound

ReflectMT: Internalizing Reflection for Efficient and High-Quality Machine Translation cites this paper.

ReflectMT: Internalizing Reflection for Efficient and High-Quality Machine Translation MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:48:27.341395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T02:45:59.473429Z digest=sha256:2bb316278897225422c470bcafb9015b097608e0708b3e6855f8bc7b90d1b4d1

Observation 5eee0f55-d720-4fea-b851-e1af78d24fd6 · inbound

JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors cites this paper.

JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.404546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T17:47:11.338622Z digest=sha256:1c5614806dd0cb211100a439b5e9dcd18bc3e173df64237a6a4a923a6b950d9f

Observation f7425033-3a6d-401a-ad10-585a4367b12a · inbound

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators cites this paper.

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T21:41:18.596975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T21:39:26.268337Z digest=sha256:38c3c03e43702cf5b7c5034120a460c235657ffb574d3dd8503744647c5bd337

Observation 466da1dc-1a78-458c-ac5b-b9df4a8e6a57 · inbound

Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages cites this paper.

Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models

Reference 195

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.718799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-03T14:31:49.214297Z digest=sha256:357d48ec4527c1af5d62a86568d2a109dc03964f6aab99a65d5a63798cdd379b