Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T13:07:22.583812Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 10 inbound Pith citation observations for arXiv:2502.07346.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T13:07:22.583812Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:16:24.014998Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
24 of 24 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 6bad7a07-f3d3-4276-8e34-4d4720866751 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 070528b1-f937-4358-8be3-459695ddc114 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aac82a55-4720-4532-8a7b-b42171e1d376 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 624aaf0c-b0ee-462c-9093-398c728096ba · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b88d08e-aafc-422f-87df-cdd1708cbdf9 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models My final verdict is tie: [[A =B]]
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a4d42ceb-2b0b-4e85-b988-23281d522e8a · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Matthew G
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d18993a2-f9e5-4663-85bb-420db724b482 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models A Survey of Neural Code Intelligence: Paradigms, Advances and Beyond
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bd5eac4-15cb-40af-ba25-0d55dc898555 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models • Math Reasoning: We collect data from MGSM which evaluates the capability of LLM to solve math reasoning problems in multiple languages, focusing on grade-school level complexity
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 35f4b71d-a579-4e97-b9de-8cf520d6aae8 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models No meaning preserved
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d45a9a21-cc8a-4118-87f9-a5e094932311 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 834550b0-d84f-47fa-8369-4d6fb52ec470 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6a3ea05d-745a-4c80-9b3e-83a085129ae4 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Only translate content in comments
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ff84a29d-7a30-4665-b823-c939508723bb · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation da4ae30e-9796-48d7-be0e-5c2e439d272b · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models [User Message] {problem} Table 18: Prompt for translating the Function Completion task
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 284da523-f5ba-415f-bd1d-927072ef8a45 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4c444337-4224-458a-89da-5fec57dc72e1 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7a4b0792-0b94-4a46-9023-9d2442ce5158 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a13c286c-75c9-4541-bd1c-bfe471a1cc43 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models [User Message] {problem} Table 19: Prompt for translating the Problem Solving task
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c8c43b53-0973-4521-9224-e299b6f2c310 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Unresolved cited work
Reference 2006
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 30b32b0b-17b8-47fc-b4bd-4c6e8eb92c36 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models InternLM2 Technical Report
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ee0ca38-c572-490d-9659-5dad7169772e · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bad54a6-ee1a-4694-843c-6ebc1b9cadd5 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e30f34c-278e-4466-bfaa-540fc3e6d816 · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb3e6816-1878-451c-8c04-f8f4646b7cda · outbound
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models DeepSeek-V3 Technical Report
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b3083bb-29df-47a1-8522-a90132a8fb6a · inbound
LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1ce3ef5-add6-418e-a613-e2fe26eca464 · inbound
MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b06a8b-5a94-403e-9b1d-3900cf7dfc19 · inbound
English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c69e8098-2699-4e5d-bd2d-87dc53a3bfe6 · inbound
Language as a Latent Variable for Reasoning Optimization BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2d81cf49-6b9d-4525-941f-e5faccb86241 · inbound
A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$\Delta$ Integration into Upcycled MoE BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation be70c081-7c96-4f4a-966b-65516e0725b4 · inbound
XLGoBench: Detecting cross-lingual skill gaps with algorithmic tasks BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7c491e6e-40ca-4251-bd7a-d8ca2df05dbf · inbound
Cross-lingual Self-Consistency for Multilingual Reasoning with Language Models BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a1f31fe5-d7ee-453f-811e-85a161d2a055 · inbound
The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 07c97628-a356-420c-8293-a32ed2ce7af5 · inbound
YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2c3cb28e-25ab-48d7-9d20-b437b9268347 · inbound
On-Policy Delta Distillation for Multilingual Math Reasoning BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.