Pith. sign in

Paper Citation Record · LEDGER

Disentangling Language and Culture for Evaluating Multilingual Large Language Models

As of 14 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2505.24635.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24635 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:33.032518Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b584db00-d352-45d2-8f3e-d42e8cb31f84 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.310743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.310743Z digest=sha256:3d4a81562049f31e1f363f882ca8d77a8f07e3c04153d20b9d1115c6234ccf98

Observation fe36fe63-bf95-4b02-aba1-fe168ef7da96 · outbound

This paper cites Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.383031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.383031Z digest=sha256:6c0136415f6ca74825061290512881dac35af5018ed886ab37c4d7b4705a3381

Observation 545853b1-e495-4d69-a88b-671cfb2c1227 · outbound

This paper cites Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.450813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.450813Z digest=sha256:f9cc7492fb2bc1f266dd081bd6e76699c4a55be5417a86f32274657a7302e0a9

Observation d6e1a612-e347-49a3-b5e8-66b55159d98f · outbound

This paper cites CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.577676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.577676Z digest=sha256:3206b0dfaed8bd8a01686872cca10e2d4eba553dd050907f4d6f7d241d675277

Observation 0f093d86-4add-44d5-b7bc-6130c22afe9c · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.670996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.670996Z digest=sha256:fd4d31f1b5bd1278307bec8ee0e917008553639a1d82b5211d5add50b40378ee

Observation e5a57296-a27f-417e-84a0-0d78ab4a73cf · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:34.318376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:35:30.786039Z digest=sha256:845aa25c3fc4f097e544305921f62fb3fb2e6c2957bbb57164025a5689068ce5

Observation ab413b52-840a-4d12-bef7-4b465ede556a · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.868536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.868536Z digest=sha256:2a46f64292df98b188efab464ba22ddf43c129f76b327fa1babf5fda460afa66

Observation 97426cd6-c677-4952-b9da-1efedfbd4f00 · outbound

This paper cites The Llama 3 Herd of Models.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.975371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.975371Z digest=sha256:cca68d3ae573c271b0b563faf57efd4a84a1859f77c79fee15b6999375d74fe1

Observation 5cc99fa0-6ce5-4cf3-bf32-eaa45ada50ad · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.063098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.063098Z digest=sha256:2f4d394aafd1a8b1cd210689dc03e3fc57d968b9a1e9a176ad973f1c4a13f8de

Observation 4c34f7a2-6192-4306-b09f-bd0a425a6909 · outbound

This paper cites Intrinsic Test of Unlearning Using Parametric Knowledge Traces.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Intrinsic Test of Unlearning Using Parametric Knowledge Traces

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.131951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.131951Z digest=sha256:1b27b8f3eed6188235a505f398cf7abc3d29d622a9d17c63a956467ecf5ba7e7

Observation a4467368-c0c4-45b9-9801-02a593adf7b1 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:34.087343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:35:31.215879Z digest=sha256:61892c53588569a3febc9cd0a1b422ccdc55ad7bf120d7a65c901433470b0933

Observation a11b7aff-1b05-483f-b640-8515b7c6c8e7 · outbound

This paper cites BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.299955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.299955Z digest=sha256:d1b90efb7dade2c978ba2e3f09e685f4122a0c8f20442a66efb118b73f4193df

Observation db639b68-38c1-4584-87dd-41e5125fb336 · outbound

This paper cites SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.361951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.361951Z digest=sha256:ba192ed6b7830e4fafbc31f529ac2cbfe7f4a71fa3c642db84c19a67c76c4707

Observation 417313c1-cc50-42f6-b690-1c93f4b9ea54 · outbound

This paper cites Crosslingual Generalization through Multitask Finetuning.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Crosslingual Generalization through Multitask Finetuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.446833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.446833Z digest=sha256:b6ddae8d6acfc9ddb8f9c7af78abd043490853daa070250a50f0ebfbbf55e776

Observation 7a09cff2-1343-4168-b394-1863624cde60 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:33.937127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:35:31.544759Z digest=sha256:c27d2cb8d707eaf77a316c36511ff6e0a847865df71795abdda27f46fcec0756

Observation b6627754-7a28-409a-a819-fe97787bc1b0 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.621695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.621695Z digest=sha256:c057648828e7fef8ba7ae3874b85fe1f1f05fb9c8adfeb4e7f054c358b757b82

Observation 5e829491-d525-4fe6-9ff4-02dc8b388965 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:33.784466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:35:31.672619Z digest=sha256:75c424a637340f10261ec74c20d287677dcef6d11ac16729df02ec39d7cfdf35

Observation 0d910d80-1073-4ae0-a48c-5777ee2eb970 · outbound

This paper cites GPT-4 Technical Report.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.753236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.753236Z digest=sha256:7f66ac1c215e58a94ffddb5e7073738615a54035c256d86f8684aeb1df3f83cc

Observation 274dd890-6f2b-4fe5-abfa-88d2fc3bf5c4 · outbound

This paper cites Qwen2.5 Technical Report.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Qwen2.5 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.893975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.893975Z digest=sha256:9a3cbfa720f83dec956c22c3afbbf4f707641d1f4f64719f4121a4daa533d188

Observation 1740e2e0-9e3d-4055-8641-d8456358748a · outbound

This paper cites Language Models are Multilingual Chain-of-Thought Reasoners.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Language Models are Multilingual Chain-of-Thought Reasoners

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.015576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.015576Z digest=sha256:4f102cfb67447a93f3b6a33216be8222c1c93807e1d6099fc7f7f0490dd0c1ec

Observation 3e978c99-7ad8-4c14-b83a-5dc943c536eb · outbound

This paper cites Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.258791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.258791Z digest=sha256:b8c6a812a4f85c4966f0cd74c04cc6ee3a2ad9f86ec92bb64599143a628eec8c

Observation 9e974ea5-6772-40b4-af04-873cf1c17ffc · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Gemma 2: Improving Open Language Models at a Practical Size

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.331555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.331555Z digest=sha256:0f014c21678def9695d3a13aad702e668040686780409c3da1490e07d3707b35

Observation 248d82b0-cc93-41e4-b373-26ab7dfe7199 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:33.597931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:35:32.434992Z digest=sha256:d991e936e14ceb0ca35779af5f3d985ae20f7554249973022b2a0c375758fb64

Observation 34fba750-ab18-4cf6-8e74-9eccf946f6b6 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.530994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.530994Z digest=sha256:4749da9d0a0c7bc2bc9209b72e65c61621459cffdfe5522eb33dd6ee0e0305ae

Observation 2d4d9e63-387f-4f55-9ef2-9ed1152d6fc5 · outbound

This paper cites Do Llamas Work in English? On the Latent Language of Multilingual Transformers.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Do Llamas Work in English? On the Latent Language of Multilingual Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.626946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.626946Z digest=sha256:031b469fbbea1f05e84dd4a18bb4f4abcf167046e516a77eeab09f2da11d2935

Observation eb7bcc25-5d00-407a-95a0-8b0f2b563f27 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.733838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.733838Z digest=sha256:918916e9d06af07ba50e5959a5b41c46ca959fdbaa1e0b9311c4406bde6af607

Observation 5dd1b83d-7f22-4f67-b534-26abce526a3c · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:33.452076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:35:32.804553Z digest=sha256:4dbaf07ad9146febdcba2f15a66b902ae6162622cc03d495d63924a1425fd409

Observation 6e5cdfa0-728f-4af4-bc80-471c884a1b90 · outbound

This paper cites SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.877616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.877616Z digest=sha256:725fbd4a53e282410083899366df4ab543b2e044e071a5a03cc4dc6804dcf19c

Observation 11197da8-c143-4207-adcd-a3ab214724d4 · outbound

This paper cites How do Large Language Models Handle Multilingualism?.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models How do Large Language Models Handle Multilingualism?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.951876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.951876Z digest=sha256:8fb35724e7c569e2ce154c3f91e5e72565b75bc239e015154827cd62c18035b0

Observation ae484c60-081a-429b-bece-63d9e534e71d · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:33.032518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:33.032518Z digest=sha256:88a309647a992580193fdd06d145f510aeaafdcec2a91a994aaa2ec4b8a0ecb5

Pith citing papers

No inbound Pith citation observations are available.