Pith. sign in

Paper Citation Record · LEDGER

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages

As of 21 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 5 inbound Pith citation observations for arXiv:2506.19468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19468 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:36:35.996643Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:17:10.649943Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved8
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 98b03044-3a20-4beb-b15a-d62e21f70207 · outbound

This paper cites https://www.anthropic.com/news/claude-3-7- sonnet.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages https://www.anthropic.com/news/claude-3-7- sonnet

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.590000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.798449Z digest=sha256:da1363719d9d81953504719a5b0bef081ed950c28e5a746f94ecef584bcfe57e

Observation 1d06bff1-adb1-4c2a-b536-8652fb2e3d64 · outbound

This paper cites https://openai.com/index/hello-gpt-4o/.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages https://openai.com/index/hello-gpt-4o/

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.575647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.805740Z digest=sha256:ad33cf6149adba1fe49b0609f5ae07a40f217d5e9863f49120fc3aa2178730a4

Observation e170ba35-2370-4d50-a6e5-7ee71e86fa26 · outbound

This paper cites https://docs.anthropic.com/en/docs/build-with-claude/multilingual- support.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages https://docs.anthropic.com/en/docs/build-with-claude/multilingual- support

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.561102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.810385Z digest=sha256:e7321f63912d2267a7db4ff470f7ab30c6c9db9c242e274c18b62f23c3692091

Observation 43d641cb-caf4-4ffe-a21d-ee96087d40b8 · outbound

This paper cites https://qwenlm.github.io/blog/qwen3/.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages https://qwenlm.github.io/blog/qwen3/

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.548739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.815557Z digest=sha256:e0863564831f65c26aa8d1a52267a6be1e474138f35b724e03ace1c41aa3c17f

Observation f2f3377b-0996-43bf-9a82-71b2c8a32d5b · outbound

This paper cites https://www.cerebras.ai/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of- redpajama.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages https://www.cerebras.ai/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of- redpajama

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.537022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.821291Z digest=sha256:e5aab3263a88846f91b51dc69c5484e7f85e026f22c6a66eb3ff05c019a4f5d1

Observation 2b33e47f-0555-4ddd-b267-a19c28153346 · outbound

This paper cites an unresolved cited work.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:36:36.524255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.826438Z digest=sha256:973278bc962c87072e5f0577c7c24f538dc78f7c3580c7629693014d56067b12

Observation fa721713-73e3-40e9-8e7f-58e2af6ec58a · outbound

This paper cites When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.508957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.831683Z digest=sha256:ac13a71da02b4a7f8f28d8c5a26134e1c56349efa773e0acf4612fbbdac49751

Observation ec99d56b-849a-4789-8fa6-08cc4a8293a5 · outbound

This paper cites Bowman, Gabor Angeli, Christopher Potts, and Christopher D.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Bowman, Gabor Angeli, Christopher Potts, and Christopher D

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T18:36:35.836435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:36:35.836435Z digest=sha256:e29c00f2392a69188f2510a6a237189f8a4b9c8fa5538704e19f027283d928cc

Observation 9bd3a764-b1b7-4488-bc5e-ca2cec99fe8d · outbound

This paper cites Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models, March 2025.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models, March 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.482994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.840727Z digest=sha256:a6c7ed720814b54fe33913b89737ede0ba3b6cec72e69109ec962f1e8c2af3e7

Observation b85ca9f9-e540-43a0-8c2d-7041d9e25d95 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge, March 2018.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge, March 2018

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.469955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.846889Z digest=sha256:d95c49f1e1952e603772709bbe3b5a38559aeb96766441dee1243dbd2a4f7667

Observation a29e9fc7-b056-4f24-b63f-28f04e968567 · outbound

This paper cites Emerging Cross-lingual Structure in Pretrained Language Models.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Emerging Cross-lingual Structure in Pretrained Language Models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.456891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.852194Z digest=sha256:7e6a35d2f57c258a19639b3914b9cd84269701bc89987a61929c3f9902b5694e

Observation f04a8942-0535-47b0-aa16-f701872a2cd8 · outbound

This paper cites Tran, Mike Zhang, Shiqi Chen, Tianyu Pang, Chao Du, Xinyi Wan, Wei Lu, and Min Lin.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Tran, Mike Zhang, Shiqi Chen, Tianyu Pang, Chao Du, Xinyi Wan, Wei Lu, and Min Lin

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.442754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.857111Z digest=sha256:56fbe7d188bc9e0dcbe6381a79ea8223f009334559e90c676b10cf459894d14a

Observation ab1e1922-a584-4e3d-bd47-2d85e3875240 · outbound

This paper cites Measuring Massive Multitask Language Understanding, January 2021.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Measuring Massive Multitask Language Understanding, January 2021

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:36:35.861228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:36:35.861228Z digest=sha256:ffe27079beb0de873d04a01c427cc82e5912fbe49526ddcc6cb989fbbf42b961

Observation 7e9b682f-6662-4a3a-88f3-bc2b4059c93e · outbound

This paper cites BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models, February 2025.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models, February 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.419675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.865204Z digest=sha256:9f84426733b96212432b1d3f380efef10233f65ec18bbf87877c3194bfd2b145

Observation 246c5349-1ce6-4dd4-874f-5a5cbe828057 · outbound

This paper cites Evaluating Code- Switching Translation with Large Language Models.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Evaluating Code- Switching Translation with Large Language Models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.404649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.869258Z digest=sha256:85205468241602b21b9abee3b25dcde6fea61220d425e2004eb06f128076eaf8

Observation fe10d072-5651-429b-9f32-826fae534430 · outbound

This paper cites ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.389923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.875617Z digest=sha256:38efedb909a6e3ab0a9b1f6865ac340b6e57668447fa66bb824fc5c8277e40ab

Observation 99bf655d-ca07-4eac-9bd6-792c4bd608be · outbound

This paper cites Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.377679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.880406Z digest=sha256:ff2785c5c6d94b5a74f92b707f1d9a4822f96f81d76565fb6dba75974f747151

Observation 23f05cb1-4611-48e0-aab3-35eaf884a8f7 · outbound

This paper cites Cross-lingual Language Model Pretraining, January 2019.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Cross-lingual Language Model Pretraining, January 2019

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.365539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.884827Z digest=sha256:4f1a693df9078f2765a3321ed454ccd4915fb005dca1f8b7e40eeb4bc63db46d

Observation 25ef6dd0-ac2e-42af-8422-ef5f55bfed11 · outbound

This paper cites The Winograd Schema Challenge.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages The Winograd Schema Challenge

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.352897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.890993Z digest=sha256:edf7d5f896b92b0c0ab74fc5c40d878223ac6bb1b41d86d875ecf444059e4643

Observation c65a92f0-4b1e-4118-83e8-6240f5cc6671 · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese, January 2024.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages CMMLU: Measuring massive multitask language understanding in Chinese, January 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.339722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.903785Z digest=sha256:3d6c22e6639a4e491b2969b330d1e4d038c9c53cd9a540a5bff500f72d27f7e1

Observation 874ebfe6-2074-419e-af6b-165b3c29fbca · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods, May 2022.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages TruthfulQA: Measuring How Models Mimic Human Falsehoods, May 2022

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.327265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.910190Z digest=sha256:ca217f55bc492e8f1b35949d50a4f1e4043f5f841af12ebce05e465462fc3c4c

Observation 548ab402-ab86-40fe-b10d-3515435425f6 · outbound

This paper cites Few-shot Learning with Multilingual Language Models, November 2022.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Few-shot Learning with Multilingual Language Models, November 2022

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.309882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.915621Z digest=sha256:ed63a6176246aa748c8f9b01e33c5ae738ce067cff0669b60988254f9e253e2e

Observation 6c13d83f-24ab-4403-b131-a2a595cd01d1 · outbound

This paper cites A Corpus and Evaluation Framework for Deeper Understanding of Commonsense Stories, April 2016.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages A Corpus and Evaluation Framework for Deeper Understanding of Commonsense Stories, April 2016

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.294228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.921919Z digest=sha256:a578d9af5d77c2b508c171c604b42a09511de5925b281d981a31cb168d890480

Observation de2a2721-bf5d-42af-8d45-bcf48bca2370 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale, October 2024.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale, October 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.282353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.926421Z digest=sha256:c9474d6db0351650b2be8e4ed56438d3689106522954e3ade47b20b122b90fee

Observation 14ab9f1e-bf65-4baf-88ce-68e172d531e5 · outbound

This paper cites Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.271759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.930581Z digest=sha256:33589810dcb0aceacdfa6721fad4d40ffb4249e63ba3e01f51e02ca2f71b0eae

Observation 741cad5f-2805-429d-bf59-a6dbd667b067 · outbound

This paper cites Qwen2.5 Technical Report, January 2025.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Qwen2.5 Technical Report, January 2025

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.259326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.934717Z digest=sha256:8f96e6d5ce164912b29ab45d4715dd676350d2e9e3ec0cb0013e176fab53c8eb

Observation 0cfd18b6-c0a0-4d3b-8a3e-e94a2ce4cd38 · outbound

This paper cites Farinha, and Alon Lavie.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Farinha, and Alon Lavie

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.249062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.938458Z digest=sha256:93ea2b385ec0ba0525bd0a0965d1e7210b27e63b9db44d9d51514bfc7bfef14f

Observation 069a478c-ba85-46ba-8347-31de6d7836c3 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T18:36:35.942564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:36:35.942564Z digest=sha256:21db299bba3a4712f02512aa7c022038e4a67b9c879496ad30007d80504b4f58

Observation 14f09d0d-5100-400d-9409-b8029b4a278b · outbound

This paper cites an unresolved cited work.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:36:36.238669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.948255Z digest=sha256:74110a8e58916ba4e7546402e3e4220958983297330e1966dba1eae4914f5770

Observation a20afbec-4b23-4406-a544-717165232b95 · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T18:36:35.953037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:36:35.953037Z digest=sha256:44d06d5be2cefd96cd920a4de1817f7e088ea8f05c3d7aa5a2076a3c8dfc7c8e

Observation 837d2353-93f2-4ac1-b566-a7bcce497b62 · outbound

This paper cites an unresolved cited work.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:36:36.226925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.959420Z digest=sha256:921539a2b3024a535eca9a2fcf5790de869a3b456d41e10a3513d983c7bf4206

Observation 3f149561-8e70-44f8-b7e7-7377b1b9afe6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models, June 2024.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Gemini: A Family of Highly Capable Multimodal Models, June 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.213558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.963708Z digest=sha256:01992bcdfb506c05df7351801f3e65ff7e884b20220e1e63858d1328c2d6d451

Observation ed20dbda-5be2-435e-a163-eb61afd3a5d1 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size, July 2024.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Gemma 2: Improving Open Language Models at a Practical Size, July 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.197316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.968257Z digest=sha256:26e9d4aeade659e0d620d23d2a6c93ccc0706bde72a1d538a56c36c5b20da9a9

Observation ac3a7651-b3d9-4f04-92a1-ff7ae8ee86f8 · outbound

This paper cites Gemma 3 Technical Report, March 2025.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Gemma 3 Technical Report, March 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.182932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.972572Z digest=sha256:7f595934070b79b7fe17be6490af2a7cad5f41a386d433c6b63c955077fbc0c1

Observation 6a2718dd-310a-4da3-9463-c30a86374681 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark, November 2024.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark, November 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.167618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.976323Z digest=sha256:76b80d3e26fa3628536590277e891d5e09b58e8a9509ea1419944983fde57ed6

Observation 01178922-80db-4fee-a5f3-525d8e439acf · outbound

This paper cites A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.153180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.980383Z digest=sha256:db4c2af93622b47c5db22a15250c17079a176324d52472edaa088619345ea9fe

Observation f531ad6d-2fd4-4cfd-9853-d90bb11031b2 · outbound

This paper cites MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation, March 2025.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation, March 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.140870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.984421Z digest=sha256:b2e8884bc81fa577c0a50ababc5cced23b33215ba833161eb5b06e4b912f5a1c

Observation a86a4da5-b507-4506-8c1b-6af4e12c5808 · outbound

This paper cites GeoM- LAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language Models.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages GeoM- LAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language Models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.124851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.988736Z digest=sha256:e2d8bb6c3471f7cd325ee1ee5b6fec75f0a2d8fb87593a4cfedae262f5d80b97

Observation e89d369c-9f7d-466d-8b69-4a16717cde6b · outbound

This paper cites an unresolved cited work.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:36:36.105733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.992934Z digest=sha256:e45334780629bf5cce1a42fbe5c8e0a5392c09280621ef36a509a0117d4419cd

Observation d708062e-4b90-4a7e-a6c4-3de2a5f2c57e · outbound

This paper cites imitative falsehoods,.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages imitative falsehoods,

Reference 40

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T18:36:36.083545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:36:35.996643Z digest=sha256:ee930ef07bbe38e244fbc47f6c5ebe689923d86633ee0de44adf3761c9caac81

Pith citing papers

Observation d7d670b6-74af-4249-82d0-ad4de16c1c5c · inbound

Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability cites this paper.

Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:00:17.818057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T20:59:59.582440Z digest=sha256:73de947ee3e2c35ce616f85df6f76117ad44e0ca3fc1939fdb014e31b6172191

Observation 730fd3d8-a5e0-4074-bf2d-508a7376e0bb · inbound

Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR cites this paper.

Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:13:15.958889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T08:04:01.791491Z digest=sha256:5d627cb26534fef8c63b64f91ba9e09df69fd1534b94da8fe2c0aa64446bb2b5

Observation 9c82c7c9-5200-4e50-86b9-ee767e5aaa4f · inbound

Creating Multilingual Mental Health Dialogue Datasets: Limits of Persona-Based Localization via Nationality and Language cites this paper.

Creating Multilingual Mental Health Dialogue Datasets: Limits of Persona-Based Localization via Nationality and Language MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T20:29:57.635151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T20:27:41.561992Z digest=sha256:478ea3d08d534dd8dff1d7a2dfe80d4cc05d5f97317b002dcab8c3bc2e0359f1

Observation 7e5a3f42-0bc6-4992-9fde-7ea61cbf2cbc · inbound

MultiHashFormer: Hash-based Generative Language Models cites this paper.

MultiHashFormer: Hash-based Generative Language Models MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T04:23:05.535623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T04:13:05.082903Z digest=sha256:a445da461f15e1c2a774037b60d1f02bdf78942561c1bd3d3ee8e2e37c96c400

Observation 9720bdf5-582d-4673-9979-f8b6660a096c · inbound

Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+ cites this paper.

Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+ MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T00:17:10.649943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:17:10.649943Z digest=sha256:7ddf253af9350121ab0401bbdafb338cab54cae797d4fe6ecfe63db6f625ae0d