Pith. sign in

Paper Citation Record · LEDGER

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia

As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2502.06298.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06298 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:57:53.734445Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:31.361951Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T10:14:36.195938Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1fba6bad-c1d8-415b-95bb-b4788f6e0c73 · outbound

This paper cites Sailor: Open Language Models for South-East Asia.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Sailor: Open Language Models for South-East Asia

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.653020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.653020Z digest=sha256:1cc8296de58d9f2aefd8cadf5498d9b76cc62af8ed4e528b3f9dd15b1a280500

Observation d2201c71-6070-48e5-81bd-6d0c5e007e1e · outbound

This paper cites The Llama 3 Herd of Models.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.656975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.656975Z digest=sha256:d8c1384dbf6af3d97d56df49453736f457e24f96e73e686c9d9015200fb06be6

Observation d06556c5-f0d9-43fc-86bf-36d016ae4cf4 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.660916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.660916Z digest=sha256:4cedf02f7a0e9de5619932c4a7ee5082f9838d6ee29c799c3ad5c9de77257717

Observation 4d5f4659-5a3c-4e58-9cb3-638d3a91b0a4 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.664955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.664955Z digest=sha256:e8f67029cc555f7973554ef1b1c179ce1fbd7180b3e81bb1f7a6480df6bcd528

Observation 84945c91-ddf1-4e28-95c0-2ee6e97be30a · outbound

This paper cites Not All Languages Are Created Equal in LLMs: Improving Multilingual Capability by Cross-Lingual-Thought Prompting.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Not All Languages Are Created Equal in LLMs: Improving Multilingual Capability by Cross-Lingual-Thought Prompting

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.672568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.672568Z digest=sha256:ea1144e3179528dc11b520d30d37b35a0c1b95f2f40d9435f0a4e1a17d305780

Observation 1d7ddba1-1f0f-4e6c-a15c-c4e34e35f74c · outbound

This paper cites A Survey on Large Language Models with Multilingualism: Recent Advances and New Frontiers.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia A Survey on Large Language Models with Multilingualism: Recent Advances and New Frontiers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.676634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.676634Z digest=sha256:af43336ee168756bc6e07d19c2cba9838431ebe5222cd0af599b1f15c454e0fd

Observation 4233678b-572b-407e-8669-689c803917de · outbound

This paper cites Mistral 7B.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Mistral 7B

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.680596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.680596Z digest=sha256:0e2918fcc6ec97759c8debe71cdc6c30877d9abfc05433b459f20c67a181c1ee

Observation 2be21d67-ecab-4335-9840-9de37aea88f2 · outbound

This paper cites ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.683952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.683952Z digest=sha256:1b79dab2254a2a82b8670ca3163e095ec250d602fa06afae3cd53b1a7247a7ef

Observation 6b66c004-bba1-4d15-a1f2-a93348aa2b50 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.687493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.687493Z digest=sha256:0d89d241f9d37e8141916e0cc57ba8f09dfff2ae95edb1c13251f798f3b105d4

Observation 066afe2d-7b7e-4bfc-a3a2-bfe5989660d5 · outbound

This paper cites WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.691083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.691083Z digest=sha256:113c1f3028b2f7895dbc9474bd8c4a2f3b99033148b223c75bd9df9db8b371ea

Observation cec73c6e-b2f7-45f6-92da-f506fc3fd98e · outbound

This paper cites Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.694595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.694595Z digest=sha256:3f692a33c52c77dadf139d31c300aa413feb7cb83a55cc58bbfac9a0ba1ab9bf

Observation 31e23f35-18db-4354-8f4f-4901da865df8 · outbound

This paper cites SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.698083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.698083Z digest=sha256:27483f5b3c5f842d6aa608b58bcebd60d1f27a78ca749c2cf23828574d6fb4aa

Observation d4e32805-fc9e-41be-877a-a010c365033d · outbound

This paper cites SeaLLMs -- Large Language Models for Southeast Asia.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia SeaLLMs -- Large Language Models for Southeast Asia

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.701632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.701632Z digest=sha256:9cb7e15365dce9c3f1d008ac7c8c69662c148d0f8cb8a1ab0499e3ea0ec81d1d

Observation f17dd090-94b8-4591-95fe-8829167c72fb · outbound

This paper cites GPT-4 Technical Report.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.705379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.705379Z digest=sha256:1e3810f20324aea70bec279b5165c4281360ce1612e9de8d1be5549f911f9414

Observation dc4e8bb6-94e1-4ea8-b9ad-52d80770ca42 · outbound

This paper cites Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.709076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.709076Z digest=sha256:092bcf7155eda3450a5d9caccdd23e67d5d72beb2abb33329e6d40fe3ee1edad

Observation 71a1e99b-ab90-4460-a9d6-258414d52ecf · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Gemma 2: Improving Open Language Models at a Practical Size

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.716047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.716047Z digest=sha256:0470e74314abf61c3a31e26b20b20489119053e06f5a1aa5f3d9f6c20e6a6325

Observation 126e4584-5d4f-46d2-b2ab-5b5ecc7e9029 · outbound

This paper cites SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural Reasoning.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.719725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.719725Z digest=sha256:dbee9b5e168ae13697e0cfd5b2e11ab8368e0c3d72a01a9940eea4b89b3bbf17

Observation 811a771a-b6f4-4a47-ab5e-0ef39f32182e · outbound

This paper cites Qwen2 Technical Report.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Qwen2 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.723180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.723180Z digest=sha256:694c81f7fab9c15678735aefa70f6092c0e6b227d8a2ac3a433b4311b06a1e86

Observation 2dd5f617-5e2c-459c-a56f-1ba1820be171 · outbound

This paper cites M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.726650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.726650Z digest=sha256:44d0bb4766290c18798c6f8d158a138c5f0c934ac47bcc78244ac9b37f6fc854

Observation b1265146-5025-4a62-9ade-7d01bbc92b3d · outbound

This paper cites SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.730210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.730210Z digest=sha256:9383905db94b87b4291f989a5529473509355cfff94090d63aed6fdcbfdb4b68

Observation 5b32dc12-74e6-4fa0-a14c-4bbdb6dde627 · outbound

This paper cites The categorization follows the practice in M3Exam (Zhang et al., 2023).

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia The categorization follows the practice in M3Exam (Zhang et al., 2023)

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:57:53.986362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:57:53.734445Z digest=sha256:d5e5e085aaaa347bfd9a116b0cf4ab4455edddd10fe806609af1fc40d70caffc

Observation d3cf97c9-a3d1-436d-bc0b-29966329a1a4 · outbound

This paper cites In Proceedings of the 2018 Conference on Empirical Methods in Nat- ural Language Processing, pages 2475–2485, Brus- sels, Belgium.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia In Proceedings of the 2018 Conference on Empirical Methods in Nat- ural Language Processing, pages 2475–2485, Brus- sels, Belgium

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:57:53.998075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:57:53.648649Z digest=sha256:db80bee4915b434dc7db78333c9595184bb200004986200a89e8f20cb67d9712

Observation f3654095-6c62-4dee-86dd-3e9d77811fd2 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Measuring Massive Multitask Language Understanding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.668748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.668748Z digest=sha256:9010174775fa89239b0089471d372a43a65e5dfd55d3d81ad1fe13bf59d21208

Observation 9fb5a892-f8a3-472e-8d93-226ed0e29ac1 · outbound

This paper cites Language Models are Multilingual Chain-of-Thought Reasoners.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Language Models are Multilingual Chain-of-Thought Reasoners

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.712567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.712567Z digest=sha256:198466a936c9fa86ae6010b894aba39076104223950fb031cc21c03f0516d175

Observation b89f303a-0c04-415f-8aea-e2d14b6bf51d · outbound

This paper cites In Proceedings of the 2023 Conference on Empiri- cal Methods in Natural Language Processing , pages 4232–4267, Singapore.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia In Proceedings of the 2023 Conference on Empiri- cal Methods in Natural Language Processing , pages 4232–4267, Singapore

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:57:54.008737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:57:53.639218Z digest=sha256:cf6fd83f65c39424247e14c6e21d5e67e74d3db9792ecb4b09724d1b919d4b6f

Observation cacd0365-b2e3-4aa9-a438-c87297933b29 · outbound

This paper cites Aya 23: Open Weight Releases to Further Multilingual Progress.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Aya 23: Open Weight Releases to Further Multilingual Progress

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.643963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.643963Z digest=sha256:1d49eff2eb7080c4c174f9354fb37bff2d15e526ec407c247089855e037a4682

Pith citing papers

Observation db639b68-38c1-4584-87dd-41e5125fb336 · inbound

Disentangling Language and Culture for Evaluating Multilingual Large Language Models cites this paper.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.361951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.361951Z digest=sha256:80f5d0e02e3a135b9e51b16703d4a02373765e3afd005a914e19f8074405ddff

Observation 85887f63-adf8-4e4d-875a-20fe503a91f5 · inbound

SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages cites this paper.

SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:14:36.197621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T10:12:45.090257Z digest=sha256:ca29919272198fd9326fdfa0008b59a1277af177c6991a38c1bce35d4a4d096b