Pith. sign in

Paper Citation Record · LEDGER

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia

As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2502.06298.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06298 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:57:53.734445Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:31.361951Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T10:14:36.195938Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1fba6bad-c1d8-415b-95bb-b4788f6e0c73 · outbound

This paper cites Sailor: Open Language Models for South-East Asia.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Sailor: Open Language Models for South-East Asia

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.653020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.653020Z digest=sha256:3dc7f7d1dcd2d6093c6ef4017fb0016d06ed7773f71df52081a4b708f1a835c9

Observation d2201c71-6070-48e5-81bd-6d0c5e007e1e · outbound

This paper cites The Llama 3 Herd of Models.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.656975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.656975Z digest=sha256:76422a1e45fdc4ea7209ca8f5c69478d15f87084bfceef03b9b8401363b1d3c7

Observation d06556c5-f0d9-43fc-86bf-36d016ae4cf4 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.660916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.660916Z digest=sha256:bd57368f181c992b3fc777824bdff4ce873d2f53010c6c7dc14d854d09706090

Observation 4d5f4659-5a3c-4e58-9cb3-638d3a91b0a4 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.664955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.664955Z digest=sha256:47dd801763b13b327a88e9c5783da099711e26e35eab06edc3c795c84d322a51

Observation 84945c91-ddf1-4e28-95c0-2ee6e97be30a · outbound

This paper cites Not All Languages Are Created Equal in LLMs: Improving Multilingual Capability by Cross-Lingual-Thought Prompting.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Not All Languages Are Created Equal in LLMs: Improving Multilingual Capability by Cross-Lingual-Thought Prompting

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.672568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.672568Z digest=sha256:e5de24385185b987e0132b861fb535089e61c816e6afd7e17a18f68e71e67931

Observation 1d7ddba1-1f0f-4e6c-a15c-c4e34e35f74c · outbound

This paper cites A Survey on Large Language Models with Multilingualism: Recent Advances and New Frontiers.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia A Survey on Large Language Models with Multilingualism: Recent Advances and New Frontiers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.676634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.676634Z digest=sha256:7c535a27e772507adc74db662632da3332d0195c2862fbdeebc056f501d59365

Observation 4233678b-572b-407e-8669-689c803917de · outbound

This paper cites Mistral 7B.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Mistral 7B

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.680596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.680596Z digest=sha256:96e9eb683ca3d0ca3c7580ee164cbbca0e4bdb82f9f60d6564bc13c0f168622a

Observation 2be21d67-ecab-4335-9840-9de37aea88f2 · outbound

This paper cites ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.683952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.683952Z digest=sha256:5b841b34bcf7a491d245740ccbf9681876ed6837623291e6a6cc744f43b0512b

Observation 6b66c004-bba1-4d15-a1f2-a93348aa2b50 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.687493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.687493Z digest=sha256:f86385b7ada92e8d41fe096d2188f0a0a7aa9f7081ef3713c38ac762bff20722

Observation 066afe2d-7b7e-4bfc-a3a2-bfe5989660d5 · outbound

This paper cites WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.691083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.691083Z digest=sha256:53c2fe14c6c4c6ae11ba165c4cefd7d91944abc1f0b80688e0eeef88eb6ae587

Observation cec73c6e-b2f7-45f6-92da-f506fc3fd98e · outbound

This paper cites Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.694595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.694595Z digest=sha256:638a5f16f45c064b1b8b779a1667e07abe5d1d3ab5afee032d7dfc707dc7d61e

Observation 31e23f35-18db-4354-8f4f-4901da865df8 · outbound

This paper cites SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.698083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.698083Z digest=sha256:d52266e63451db19865ce68c9a80e281b9e0e4052b36627975dfa4a7e14b8953

Observation d4e32805-fc9e-41be-877a-a010c365033d · outbound

This paper cites SeaLLMs -- Large Language Models for Southeast Asia.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia SeaLLMs -- Large Language Models for Southeast Asia

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.701632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.701632Z digest=sha256:5f7180cffb26eef280af0c4aa080b4179f66587a6e2b0b06c63f0cdfbae092c5

Observation f17dd090-94b8-4591-95fe-8829167c72fb · outbound

This paper cites GPT-4 Technical Report.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.705379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.705379Z digest=sha256:457d6b0a8a080bf41f0035c3607f4adfcff3ff66a89aa4e062a8206b77be92c2

Observation dc4e8bb6-94e1-4ea8-b9ad-52d80770ca42 · outbound

This paper cites Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.709076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.709076Z digest=sha256:859d3c20eb212d03a320c7a0a0da1ec6b3534f4472c8a57396680be4ad60b8c6

Observation 71a1e99b-ab90-4460-a9d6-258414d52ecf · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Gemma 2: Improving Open Language Models at a Practical Size

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.716047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.716047Z digest=sha256:4c522464462188ee25fa87e3f9eaf4f9ab99b8ba6c2e60ca5a906c7550dfddd2

Observation 126e4584-5d4f-46d2-b2ab-5b5ecc7e9029 · outbound

This paper cites SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural Reasoning.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.719725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.719725Z digest=sha256:acead18e50e0da31592d82b8aad21c5586bc067ede2f3d9a66e0323881f3d75c

Observation 811a771a-b6f4-4a47-ab5e-0ef39f32182e · outbound

This paper cites Qwen2 Technical Report.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Qwen2 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.723180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.723180Z digest=sha256:6bd103e238e8fb781f6819ebdf44adda0d390da169217be5eb71661b916aa88b

Observation 2dd5f617-5e2c-459c-a56f-1ba1820be171 · outbound

This paper cites M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.726650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.726650Z digest=sha256:b091440990a9eb0cb6bf84448a71ec2ef7b9055af9727307f268985e302e51fe

Observation b1265146-5025-4a62-9ade-7d01bbc92b3d · outbound

This paper cites SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.730210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.730210Z digest=sha256:5efc0a60d65a09d02ef163945fdcdf93335adca9988fda22c3b8dec3126d60a2

Observation 5b32dc12-74e6-4fa0-a14c-4bbdb6dde627 · outbound

This paper cites The categorization follows the practice in M3Exam (Zhang et al., 2023).

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia The categorization follows the practice in M3Exam (Zhang et al., 2023)

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:57:53.986362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:57:53.734445Z digest=sha256:848835104075ddd7550b1cb09c225fc5f7ed6481c41f45e9b4222495f85731c2

Observation d3cf97c9-a3d1-436d-bc0b-29966329a1a4 · outbound

This paper cites In Proceedings of the 2018 Conference on Empirical Methods in Nat- ural Language Processing, pages 2475–2485, Brus- sels, Belgium.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia In Proceedings of the 2018 Conference on Empirical Methods in Nat- ural Language Processing, pages 2475–2485, Brus- sels, Belgium

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:57:53.998075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:57:53.648649Z digest=sha256:f6723cb8cd518d981c093032a0449051a8936a386ee901c49ceb3ccba2d13260

Observation f3654095-6c62-4dee-86dd-3e9d77811fd2 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Measuring Massive Multitask Language Understanding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.668748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.668748Z digest=sha256:e2e156e6530cc1c8f93e59fb8d42817cfc83f7aaaa4dab5e1101eee5bcd6b1fe

Observation 9fb5a892-f8a3-472e-8d93-226ed0e29ac1 · outbound

This paper cites Language Models are Multilingual Chain-of-Thought Reasoners.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Language Models are Multilingual Chain-of-Thought Reasoners

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.712567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.712567Z digest=sha256:fe417e91167d5d5951968c9c4b2ce2e473a277e979d4d93e97451d97cc333078

Observation b89f303a-0c04-415f-8aea-e2d14b6bf51d · outbound

This paper cites In Proceedings of the 2023 Conference on Empiri- cal Methods in Natural Language Processing , pages 4232–4267, Singapore.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia In Proceedings of the 2023 Conference on Empiri- cal Methods in Natural Language Processing , pages 4232–4267, Singapore

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:57:54.008737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:57:53.639218Z digest=sha256:206e2cabdd5f65054fce2f106640937f2ea1d56eddde75f9cf9c1ef55e1a3673

Observation cacd0365-b2e3-4aa9-a438-c87297933b29 · outbound

This paper cites Aya 23: Open Weight Releases to Further Multilingual Progress.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Aya 23: Open Weight Releases to Further Multilingual Progress

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.643963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.643963Z digest=sha256:7f4b356fda85add9ccc16eb0fd48350169692addc6bb132d4c78177e1222f7f1

Pith citing papers

Observation db639b68-38c1-4584-87dd-41e5125fb336 · inbound

Disentangling Language and Culture for Evaluating Multilingual Large Language Models cites this paper.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.361951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.361951Z digest=sha256:c4865fcf71ec731917f93a04d5c381df384fc22d14679366ae95693be43d107b

Observation 85887f63-adf8-4e4d-875a-20fe503a91f5 · inbound

SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages cites this paper.

SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:14:36.197621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T10:12:45.090257Z digest=sha256:dac8e71aa13ef2270cb142acaa0cb3206098b033015355705b073713e5b91a0b