Pith. sign in

Paper Citation Record · LEDGER

Disentangling Language and Culture for Evaluating Multilingual Large Language Models

As of 8 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2505.24635.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24635 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:33.032518Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b584db00-d352-45d2-8f3e-d42e8cb31f84 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.310743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.310743Z digest=sha256:3ccda003e70e1571766fca6dab14f59949bebc844a6680bc1f432f2016453fea

Observation fe36fe63-bf95-4b02-aba1-fe168ef7da96 · outbound

This paper cites Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.383031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.383031Z digest=sha256:39cb75173ea41aa41adbd56ef3424dbb86ac3e35a9c1dfa01b49636f55913c3e

Observation 545853b1-e495-4d69-a88b-671cfb2c1227 · outbound

This paper cites Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.450813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.450813Z digest=sha256:78bf89f31c2ac3555d4fc10bfe38fe5b64f49f1c270d9598fa647eac281a76a4

Observation d6e1a612-e347-49a3-b5e8-66b55159d98f · outbound

This paper cites CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.577676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.577676Z digest=sha256:3fc1dae9b49533567cbbef2abb933674cb90667f4d1ae03000e2e5bcca8e063b

Observation 0f093d86-4add-44d5-b7bc-6130c22afe9c · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.670996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.670996Z digest=sha256:cb5e929ec5b2e88e164fbed40876dc754c9334e19ba10f679b5460b3a889fedc

Observation e5a57296-a27f-417e-84a0-0d78ab4a73cf · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:34.318376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:35:30.786039Z digest=sha256:98797597399b109dfa02ee203804a88483b0e1e128a97d218659023e3dd0c60b

Observation ab413b52-840a-4d12-bef7-4b465ede556a · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.868536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.868536Z digest=sha256:bd2c37234eac9c105b0f7a13c334c2da9fe2b49a384c4e902b19c6a23daba49e

Observation 97426cd6-c677-4952-b9da-1efedfbd4f00 · outbound

This paper cites The Llama 3 Herd of Models.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.975371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.975371Z digest=sha256:45fc668cb8014084f814a90e4bbed72dee2b8a9f596327f705e50168703f2929

Observation 5cc99fa0-6ce5-4cf3-bf32-eaa45ada50ad · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.063098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.063098Z digest=sha256:d12efd703f5d635270ef507fdfc486d8f8001cc5761bda2ee380204a5ab41d63

Observation 4c34f7a2-6192-4306-b09f-bd0a425a6909 · outbound

This paper cites Intrinsic Test of Unlearning Using Parametric Knowledge Traces.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Intrinsic Test of Unlearning Using Parametric Knowledge Traces

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.131951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.131951Z digest=sha256:1e54468713db23ea67716cc5ca4367752645edcb159771cdedfc101cfc506392

Observation a4467368-c0c4-45b9-9801-02a593adf7b1 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:34.087343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:35:31.215879Z digest=sha256:459d0089065d1e6f11aec85c6d157eb51c9932913f09a9b7387e2d24be442759

Observation a11b7aff-1b05-483f-b640-8515b7c6c8e7 · outbound

This paper cites BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.299955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.299955Z digest=sha256:2ba8404f0788ed91f37977b70624acf3eb351adcc4d1e1dee0cf4d6f6b0d87a6

Observation db639b68-38c1-4584-87dd-41e5125fb336 · outbound

This paper cites SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.361951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.361951Z digest=sha256:c17aa924b62dd0f7a2b31970d3d5081bffa6551b3df7617703ad01ea90288609

Observation 417313c1-cc50-42f6-b690-1c93f4b9ea54 · outbound

This paper cites Crosslingual Generalization through Multitask Finetuning.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Crosslingual Generalization through Multitask Finetuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.446833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.446833Z digest=sha256:ff7e13f6896d8af77a02fbec3a583c058f9aad64b6def0497f4a427800a68185

Observation 7a09cff2-1343-4168-b394-1863624cde60 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:33.937127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:35:31.544759Z digest=sha256:05ef0c2260d332f675d859296cba1da5ff84967531199d050426a8d9cd042c54

Observation b6627754-7a28-409a-a819-fe97787bc1b0 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.621695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.621695Z digest=sha256:639a5df33b1cc52144e1bb54ef33b3ccbca6e541a11b4f1b025bff26731b5363

Observation 5e829491-d525-4fe6-9ff4-02dc8b388965 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:33.784466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:35:31.672619Z digest=sha256:0c3e1bc94aed104abe0ab4a6f76d4ef978e192ec84867990e82dd6d4f35697bb

Observation 0d910d80-1073-4ae0-a48c-5777ee2eb970 · outbound

This paper cites GPT-4 Technical Report.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.753236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.753236Z digest=sha256:46d9123680e1530c19ebc7cb8bad291dfd591450d57a5a49d9d8cea153239405

Observation 274dd890-6f2b-4fe5-abfa-88d2fc3bf5c4 · outbound

This paper cites Qwen2.5 Technical Report.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Qwen2.5 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.893975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.893975Z digest=sha256:35bfd425bd3cd3748a57da4da081c3169f63fc4c72fa85bbeca085c467b8d0b0

Observation 1740e2e0-9e3d-4055-8641-d8456358748a · outbound

This paper cites Language Models are Multilingual Chain-of-Thought Reasoners.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Language Models are Multilingual Chain-of-Thought Reasoners

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.015576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.015576Z digest=sha256:69f2784c9b4f6aab53ac6eb85ccef82cc75d136cf0b8fe526b38348191544429

Observation 3e978c99-7ad8-4c14-b83a-5dc943c536eb · outbound

This paper cites Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.258791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.258791Z digest=sha256:3514ff4dab6264f431f8ce5914b38074f259acd82538ee2a03c2a6b999d9c924

Observation 9e974ea5-6772-40b4-af04-873cf1c17ffc · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Gemma 2: Improving Open Language Models at a Practical Size

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.331555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.331555Z digest=sha256:c1dbe87e741eafa54df66deb8d4156a59a361f70dcc58d465f917551bef860ed

Observation 248d82b0-cc93-41e4-b373-26ab7dfe7199 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:33.597931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:35:32.434992Z digest=sha256:26ac83ac3fd8d66c53fb0b3db065a4aee2cee13765bb9632f812d57498f9559c

Observation 34fba750-ab18-4cf6-8e74-9eccf946f6b6 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.530994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.530994Z digest=sha256:611bc678f439fb0dd9d98c3ac2a919fbd93afd4e272cb2cbbd38a3d50d3a2075

Observation 2d4d9e63-387f-4f55-9ef2-9ed1152d6fc5 · outbound

This paper cites Do Llamas Work in English? On the Latent Language of Multilingual Transformers.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Do Llamas Work in English? On the Latent Language of Multilingual Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.626946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.626946Z digest=sha256:f818525daf767e8440d7b61f430ffbc02c27551a16b1dc2e9c12063bf17b7322

Observation eb7bcc25-5d00-407a-95a0-8b0f2b563f27 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.733838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.733838Z digest=sha256:050e7ef8a1c0cc72e389461c84be7c15902d213c84df29402b7387d7aa77fa87

Observation 5dd1b83d-7f22-4f67-b534-26abce526a3c · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:33.452076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:35:32.804553Z digest=sha256:564093d3881f7ff43451b80195ac6d26f30fc7d4305bcf72b3a8229129ce4501

Observation 6e5cdfa0-728f-4af4-bc80-471c884a1b90 · outbound

This paper cites SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.877616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.877616Z digest=sha256:5b021e9431e904e3da666429e61070c8b6bc51dec3dff4dc81acc475eee457a5

Observation 11197da8-c143-4207-adcd-a3ab214724d4 · outbound

This paper cites How do Large Language Models Handle Multilingualism?.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models How do Large Language Models Handle Multilingualism?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.951876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.951876Z digest=sha256:c71709088e8a11cabb29dd9b9315a50e619aacea83e436d5884c454a00b7e329

Observation ae484c60-081a-429b-bece-63d9e534e71d · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:33.032518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:33.032518Z digest=sha256:fcf2f86c3edded899f0315f92f52d99c022efb7d32b4fbc2b35ae6cec1c9af33

Pith citing papers

No inbound Pith citation observations are available.